Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your path

Q50HardConcept

Walk through everything that happens from the moment a user sends a prompt to a hosted LLM until the response streams back.

30-second answerSay your answer out loud first, then reveal.
A sequence chart of one request: user to gateway to scheduler, prefill on the model GPUs, a decode loop that streams each token back to the user, then usage and the finish reason returned through the gateway.

Step-by-step detail

  1. Gateway: API key / auth, quota and rate limits, request validation, possibly input moderation.
  2. Routing: pick a replica based on model, load and prefix-cache affinity.
  3. Preprocessing: apply the chat template (system/user/assistant markers, tool definitions), tokenize into IDs, check context length.
  4. Scheduling: the request joins the continuous batch; KV blocks are allocated (PagedAttention); cached prefix blocks are reused if available.
  5. Prefill: embeddings → N transformer layers (attention + MLP, distributed across GPUs with tensor parallelism) → KV cache for all prompt tokens → logits for the next token. Determines TTFT.
  6. Decode loop:
    Logits → temperature/top-p, penalties, logit bias, grammar mask (structured output) → sample a token.
    Detokenize incrementally (handling multi-byte characters) and stream via Server-Sent Events.
    Stop on EOS, stop sequence, max tokens, or tool-call completion.
  7. Post-processing: output safety filtering (streaming-compatible), tool-call parsing, usage accounting (input/output/cached tokens), logging and metrics (TTFT, TPOT).
Why interviewers ask this. It checks that you can connect concepts (tokenization, attention, KV cache, batching, sampling, serving) into one coherent system view. That's a strong signal for senior AI engineering roles.

Every expert started right here.