Choosing a model
A model is a trained neural network, named by an id such as openai/gpt-oss-120b, and choosing one means trading answer quality against speed and price.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
With a hosted API you never load a model yourself: every request names the model it wants, and the provider runs it. Changing models is changing one string.
The video describes a language model by two numbers. Parameters are its weights and biases, the numbers it learned, counted in billions or trillions. Tokens are the pieces of text it was trained on, also counted in trillions. The transformer is the architecture underneath. The video then uses parameters to sort models: below about 10 billion it calls a model a small language model (SLM).
That 10 billion line is a rule of thumb from the video, not an official definition. The two models below carry their size in their names: gpt-oss-20b has about 20 billion parameters and gpt-oss-120b about 120 billion.
Syntax: client.models.list() returns the models your key can use; model= in a request picks one.
Listing the models your key can use
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
for model in client.models.list().data:
print(model.id)whisper-large-v3 openai/gpt-oss-safeguard-20b canopylabs/orpheus-v1-english meta-llama/llama-prompt-guard-2-22m canopylabs/orpheus-arabic-saudi qwen/qwen3.8-27b openai/gpt-oss-20b allam-2-7b openai/gpt-oss-120b meta-llama/llama-prompt-guard-2-86m whisper-large-v3-turbo
- Not every id is a chat model. The
whispermodels turn speech into text, theorpheusmodels turn text into speech, and theprompt-guardmodels only classify prompts. - The chat models include
openai/gpt-oss-20bandopenai/gpt-oss-120b, the pair this lesson compares. - The order changes from call to call; sort the ids if you want a stable list.
Details of the two gpt-oss models
models = {model.id: model for model in client.models.list().data}
for name in ["openai/gpt-oss-20b", "openai/gpt-oss-120b"]:
model = models[name]
print(name, "| made by", model.owned_by, "| weights:", model.hugging_face_id)openai/gpt-oss-20b | made by OpenAI | weights: openai/gpt-oss-20b openai/gpt-oss-120b | made by OpenAI | weights: openai/gpt-oss-120b
Both are made by OpenAI, and their weights are published on Hugging Face under the same names. That is what open-weight means: you could download them and run them yourself, on hardware with enough memory. Groq runs them for you.
Asking the 20b and the 120b the same question
question = [{"role": "user", "content": "In one sentence, what is a refund?"}]
for name in ["openai/gpt-oss-20b", "openai/gpt-oss-120b"]:
response = client.chat.completions.create(model=name, messages=question)
print(name)
print(" ", response.choices[0].message.content)openai/gpt-oss-20b A refund is the return of money paid for a product or service, usually given to a buyer who has returned or is dissatisfied with the purchase. openai/gpt-oss-120b A refund is the return of money to a buyer when a purchased product or service is returned, canceled, or found unsatisfactory.
What the two answers show
- Both answers are correct and one sentence long, as asked.
- The wording differs. The 20b answer mentions a buyer who is dissatisfied; the 120b answer lists three cases: returned, cancelled or unsatisfactory.
- An easy question hides the difference in size. Larger models pull ahead on harder tasks: long instructions, tricky categories, several steps of reasoning. Project: ticket sorter swaps models on a real task and measures it.
Hosted model vs open weights on your computer
| Hosted (this course) | Open weights on your computer | |
|---|---|---|
| Setup | An API key | Download the weights, install a runtime |
| Hardware | None of yours | Enough memory for the weights, often a GPU |
| Cost | Per token | Your hardware and electricity |
| Your data | Sent to the provider | Stays on your machine |
| Changing model | Change the id | Download another model |
When to pick the smaller model
- A simple, well-defined task, such as sorting tickets into three categories, where the smaller model's score is good enough.
- High volume, where half the price per token adds up (Cost reads both prices).
- When speed matters more than polish.
model_not_found error. Keep the id in one variable, and when that error appears, list the models again to find the replacement.Related
- Previous: Installation and setup
- Next: Tokens
- Reference: Groq supported models
- Print
model.context_windowfor both gpt-oss models. Context window explains the number. - Ask both models a harder question, such as a two-step word problem, and compare the answers.
- Print
sorted(model.id for model in client.models.list().data)to get the ids in a stable order.
You understood something today that you didn't yesterday.