1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you reduce an agent's latency?
30-second answerSay your answer out loud first, then reveal.
Where time goes: total latency ≈ Σ (LLM time-to-first-token + generation time) + Σ tool latency, across sequential steps. Long outputs and many steps are the usual culprits.
Techniques
- Fewer steps: combine tools so one call does what three did; hard-code predictable steps; give the agent all the context it needs upfront so it doesn't need discovery calls.
- Parallelism: parallel tool calls in one turn, parallel sub-agents, and speculative execution (start likely tool calls early).
- Model routing: a small, fast model for classification, extraction and simple tool picks; the large model only where reasoning is needed.
- Fewer output tokens: generation is the slow part. Ask for concise reasoning and structured short outputs; limit verbose chain-of-thought where it doesn't help.
- Prompt caching: stable prefixes (system prompt, tool definitions) reduce time-to-first-token on long prompts.
- Streaming / progress UX: stream the final answer, and show progress ("Searching orders...") to cut perceived latency.
- Faster tools: caching, indexes, timeouts, async I/O.
- Async mode: for long tasks, return immediately and notify the user when done.
Answer structure tip. "First I'd trace p50/p95 latency by span to find where time goes, then..." Measuring first is what interviewers listen for.
Related
You understood something today that you didn't yesterday.