1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
After a deploy, p95 latency doubled. How do you investigate?
30-second answerSay your answer out loud first, then reveal.
Investigation steps
- Mitigate first: if the SLO error budget is burning fast, roll back (flag or config), then debug.
- Diff the release bundle: what changed in code, prompt, model, retrieval or guardrails?
- Per-span comparison (before vs after):
| Observation | Likely cause |
|---|---|
| Input tokens +60% | Bigger prompt, more retrieved chunks, history handling bug |
| Output tokens +80% | Prompt change made answers verbose; max_tokens raised |
| TTFT up, same tokens | Prompt caching broken (dynamic content moved to the start); different model/region |
| Extra LLM span | New step (rewrite, judge) added sequentially |
| Retry spans | JSON validation failures with the new prompt |
| Tool spans slower | New tool or changed API usage |
4. Check external factors: provider status, traffic mix shift, a canary on different infrastructure.
5. Fix and verify with a load test or canary; add a latency regression check (tokens and p95) to the CI gates.
Related
Slow is fine. Stopping is the only problem.