1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →
Latency
Latency is the time between sending a request and getting the answer back, and for an LLM much of it is spent writing output tokens one after another.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
A person waiting for a reply notices every second. Two numbers describe the wait: how long before the first words appear, and how long the whole answer takes.
Syntax: time.perf_counter() before and after a request gives the seconds it took.
Timing a short answer and a long one
import time
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
for request in ["Reply with one word: yes", "Write about 150 words on how parcels are delivered."]:
started = time.perf_counter()
response = client.chat.completions.create(
model="openai/gpt-oss-120b", messages=[{"role": "user", "content": request}]
)
seconds = time.perf_counter() - started
print(f"{response.usage.completion_tokens} tokens in {seconds:.2f} seconds")Output
51 tokens in 0.77 seconds 246 tokens in 1.04 seconds
- 51 tokens for a one-word answer. Most of them are the model's reasoning, which is written before the word.
- 246 tokens took about a third longer, not five times longer. Groq writes hundreds of tokens a second, so the fixed part of every request, the network and the queue, is a large share of a short request.
- Your numbers will differ from run to run; compare the two lines, not the exact seconds.
Time to the first words
This continues the same file, with stream=True from Generating text:
started = time.perf_counter()
first = None
stream = client.chat.completions.create(
model="openai/gpt-oss-120b",
messages=[{"role": "user", "content": "Write about 150 words on how parcels are delivered."}],
stream=True,
)
for chunk in stream:
if first is None and chunk.choices and chunk.choices[0].delta.content:
first = time.perf_counter() - started
total = time.perf_counter() - started
if first is None:
print(f"no answer text came back after {total:.2f} seconds")
else:
print(f"first words after {first:.2f} seconds, the whole answer after {total:.2f} seconds")Output
first words after 0.60 seconds, the whole answer after 1.01 seconds
- First words after 0.60 seconds, the whole answer 0.41 seconds later. More than half of the wait came before the first word.
- The reasoning comes first. gpt-oss writes its reasoning before the answer, and those chunks carry no answer text, so the loop sees no words until the reasoning ends.
- The check for
Nonecovers a reply with no answer text, which can happen; without it the last line would fail onNone. - Times move from run to run. Run it a few times before reading anything into one number.
Time to first token vs tokens per second
| Time to first token | Tokens per second | |
|---|---|---|
| Measures | The wait before anything appears | How fast the rest follows |
| Grows with | Prompt length, reasoning, queue | Model size, provider |
| Seen by | Anyone watching a streamed answer | Anyone waiting for a long answer |
Ways to make answers faster
- Ask for shorter answers.
- Lower
reasoning_effortfor simple tasks (Max tokens and finish_reason). - Use a smaller model where it scores well enough.
- Send independent requests at the same time, as the async lessons in Python for AI do.
Watch out. Streaming does not shorten the total time. It shows the answer while it is written, which makes a wait feel shorter, but the last token arrives when it would have anyway.
Related
- Previous: Cost
- Next: Project: ticket sorter
- Reference: Groq reasoning models
Try it yourself
- Add
reasoning_effort="low"to both requests in the first example and compare the tokens and seconds. - Run the first-token example three times and note how much the times move.
- Print
response.usage.completion_timeafter the long request.
You understood something today that you didn't yesterday.