Max tokens and finish_reason
max_completion_tokens is a request setting that caps how many tokens the model may write, and finish_reason in the response says whether the answer ended by itself (stop) or was cut off at the cap (length).
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
An answer ends in one of two ways: the model writes its end token, or it runs into the cap you set. Only the first means the answer is finished.
Syntax: client.chat.completions.create(..., max_completion_tokens=300), then response.choices[0].finish_reason.
A cap of 30 tokens and a cap of 300
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
prompt = [{"role": "user", "content": "Write a one-sentence apology to a customer whose parcel is late."}]
for limit in (30, 300):
response = client.chat.completions.create(
model="openai/gpt-oss-120b", messages=prompt, max_completion_tokens=limit
)
choice = response.choices[0]
print(limit, choice.finish_reason, repr(choice.message.content))30 length '' 300 stop 'We sincerely apologize for the delay in delivering your parcel and appreciate your patience while we work to get it to you as quickly as possible.'
- 30, length, '': the answer was cut off at the cap, and the text is empty. Not one word of the apology was written.
- 300, stop: the model finished on its own, with a full sentence.
Where the 30 tokens went
response = client.chat.completions.create(
model="openai/gpt-oss-120b", messages=prompt, max_completion_tokens=300
)
message = response.choices[0].message
print("reasoning:", message.reasoning)
print("answer:", message.content)
print(response.usage.completion_tokens, "completion tokens,",
response.usage.completion_tokens_details.reasoning_tokens, "of them reasoning")reasoning: The user wants a one-sentence apology to a customer whose parcel is late. Provide a concise apology. answer: We’re sorry for the delay—your parcel is on its way and we appreciate your patience while we work to get it to you as quickly as possible. 62 completion tokens, 22 of them reasoning
gpt-oss is a reasoning model: before the answer, it writes a short note to itself, returned as message.reasoning. Those tokens count towards max_completion_tokens and towards the bill. Here 22 of the 62 completion tokens were reasoning, so a cap of 30 left almost nothing for the answer.
Lowering the reasoning effort
response = client.chat.completions.create(
model="openai/gpt-oss-120b", messages=prompt, max_completion_tokens=60, reasoning_effort="low"
)
choice = response.choices[0]
print(choice.finish_reason, repr(choice.message.content))stop 'We’re sorry for the delay in delivering your parcel and appreciate your patience while we work to get it to you as quickly as possible.'
reasoning_effort="low" asks gpt-oss for less reasoning. With a cap of 60, the answer finished with stop. The other values are "medium", the default, and "high".
stop vs length
| finish_reason | Meaning | What to do |
|---|---|---|
stop | The model ended the answer itself | Use the answer |
length | The cap cut the answer off | Raise the cap, lower the reasoning effort, or ask for a shorter answer |
When to set a cap
- To stop a runaway answer from using up time and tokens.
- To keep costs predictable when many requests run at once.
- Always with room for the reasoning, when the model is a reasoning model.
finish_reason before you use an answer. A JSON answer cut off at the cap is not valid JSON, and length tells you why before your parser fails.Related
- Previous: top_p and top_k
- Next: Seeds
- Reference: Groq reasoning models
- Set the cap to
60withoutreasoning_effortand checkfinish_reason. - Print
response.usage.completion_tokens_details.reasoning_tokenswithreasoning_effort="high". - Ask for a three-paragraph apology with a cap of 100 and read where it stops.
Slow is fine. Stopping is the only problem.