1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
A month after launch, the LLM bill is 5x the forecast. How do you investigate and fix it?
30-second answerSay your answer out loud first, then reveal.
Investigation
- Attribution: tag every request with feature, tenant, user and model. Without this you can't diagnose anything, so that's the first fix if it's missing.
- Distribution analysis: cost per request percentiles. Is it uniform growth or a long tail of extreme requests?
- Common findings:
| Finding | Fix |
|---|---|
| Input tokens per turn grow linearly with conversation length | Summarise or trim history; cap turns |
| Retrieved context of 20 chunks per query | Rerank, send top 5 |
| Agent runs averaging 25 steps | Better tools, step limits, loop detection |
| Retries on validation failures (3x calls) | Fix prompts, use structured outputs |
| Frontier model for intent classification | Route to a small model |
| Same system prompt re-sent without caching | Enable prompt caching; stable prefix |
| 1% of users generate 40% of cost (scripts/abuse) | Per-user quotas, abuse detection |
| Forecast assumed shorter outputs | Max token limits, concise style instructions |
- Quick wins vs structural fixes: caching and routing in days; architecture changes over weeks.
- Governance: budgets per feature, alerts on daily spend anomalies, cost visible on the team dashboard, cost regression checks in CI evals (tokens per test case).
Key metric: cost per successful outcome (e.g. per resolved ticket), not just total spend.
Related
Little by little, you're building something great.