1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Leadership asks you to cut LLM inference costs by 50% this quarter without hurting quality. What's your plan?
30-second answerSay your answer out loud first, then reveal.
Plan
| Week | Action | Typical impact |
|---|---|---|
| 1–2 | Cost attribution + per-request distributions; identify the top 3 features (often ~80% of spend) | Visibility |
| 2–3 | Remove waste: retry storms, agent loops, duplicated calls, abuse; cap max_tokens | 5–20% |
| 3–5 | Prompt caching (restructure prompts so the static prefix comes first); semantic/exact caches for repeated queries | 10–30% on input-heavy features |
| 4–7 | Context diet: fewer retrieved chunks with a reranker, summarised history, shorter system prompts | 10–25% |
| 5–9 | Model routing: small model for classification, extraction, simple Q&A; cascade to a large model on low confidence | 20–50% on routed traffic |
| 6–10 | Move offline jobs to batch APIs | ~50% on those jobs |
| 8–12 | Distil or fine-tune a small model for the highest-volume narrow task; consider self-hosting if volume justifies it | Large for that task |
Guardrails: eval gates for every change (quality must stay within agreed tolerance), canary rollouts, a cost-per-outcome dashboard, and a weekly progress report.
Interview signal. A quantified, measured, quality-protected plan, rather than "switch to a cheaper model".
Related
Little by little, you're building something great.