Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q27IntermediateConcept

How do you reduce an agent's latency?

30-second answerSay your answer out loud first, then reveal.

Where time goes: total latency ≈ Σ (LLM time-to-first-token + generation time) + Σ tool latency, across sequential steps. Long outputs and many steps are the usual culprits.

Techniques

  1. Fewer steps: combine tools so one call does what three did; hard-code predictable steps; give the agent all the context it needs upfront so it doesn't need discovery calls.
  2. Parallelism: parallel tool calls in one turn, parallel sub-agents, and speculative execution (start likely tool calls early).
  3. Model routing: a small, fast model for classification, extraction and simple tool picks; the large model only where reasoning is needed.
  4. Fewer output tokens: generation is the slow part. Ask for concise reasoning and structured short outputs; limit verbose chain-of-thought where it doesn't help.
  5. Prompt caching: stable prefixes (system prompt, tool definitions) reduce time-to-first-token on long prompts.
  6. Streaming / progress UX: stream the final answer, and show progress ("Searching orders...") to cut perceived latency.
  7. Faster tools: caching, indexes, timeouts, async I/O.
  8. Async mode: for long tasks, return immediately and notify the user when done.
Answer structure tip. "First I'd trace p50/p95 latency by span to find where time goes, then..." Measuring first is what interviewers listen for.

You understood something today that you didn't yesterday.