LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Latency

Latency is the time between sending a request and getting the answer back, and for an LLM much of it is spent writing output tokens one after another.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

A person waiting for a reply notices every second. Two numbers describe the wait: how long before the first words appear, and how long the whole answer takes.

Syntax: time.perf_counter() before and after a request gives the seconds it took.

Timing a short answer and a long one

ExampleAPI key
import time

from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment

for request in ["Reply with one word: yes", "Write about 150 words on how parcels are delivered."]:
    started = time.perf_counter()
    response = client.chat.completions.create(
        model="openai/gpt-oss-120b", messages=[{"role": "user", "content": request}]
    )
    seconds = time.perf_counter() - started
    print(f"{response.usage.completion_tokens} tokens in {seconds:.2f} seconds")
  • 51 tokens for a one-word answer. Most of them are the model's reasoning, which is written before the word.
  • 246 tokens took about a third longer, not five times longer. Groq writes hundreds of tokens a second, so the fixed part of every request, the network and the queue, is a large share of a short request.
  • Your numbers will differ from run to run; compare the two lines, not the exact seconds.

Time to the first words

This continues the same file, with stream=True from Generating text:

ExampleAPI key
started = time.perf_counter()
first = None
stream = client.chat.completions.create(
    model="openai/gpt-oss-120b",
    messages=[{"role": "user", "content": "Write about 150 words on how parcels are delivered."}],
    stream=True,
)
for chunk in stream:
    if first is None and chunk.choices and chunk.choices[0].delta.content:
        first = time.perf_counter() - started
total = time.perf_counter() - started
if first is None:
    print(f"no answer text came back after {total:.2f} seconds")
else:
    print(f"first words after {first:.2f} seconds, the whole answer after {total:.2f} seconds")
  • First words after 0.60 seconds, the whole answer 0.41 seconds later. More than half of the wait came before the first word.
  • The reasoning comes first. gpt-oss writes its reasoning before the answer, and those chunks carry no answer text, so the loop sees no words until the reasoning ends.
  • The check for None covers a reply with no answer text, which can happen; without it the last line would fail on None.
  • Times move from run to run. Run it a few times before reading anything into one number.

Time to first token vs tokens per second

Time to first tokenTokens per second
MeasuresThe wait before anything appearsHow fast the rest follows
Grows withPrompt length, reasoning, queueModel size, provider
Seen byAnyone watching a streamed answerAnyone waiting for a long answer

Ways to make answers faster

  • Ask for shorter answers.
  • Lower reasoning_effort for simple tasks (Max tokens and finish_reason).
  • Use a smaller model where it scores well enough.
  • Send independent requests at the same time, as the async lessons in Python for AI do.
Watch out. Streaming does not shorten the total time. It shows the answer while it is written, which makes a wait feel shorter, but the last token arrives when it would have anyway.
Try it yourself
  • Add reasoning_effort="low" to both requests in the first example and compare the tokens and seconds.
  • Run the first-token example three times and note how much the times move.
  • Print response.usage.completion_time after the long request.
PreviousCost

You understood something today that you didn't yesterday.