Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q41HardScenario

Leadership asks you to cut LLM inference costs by 50% this quarter without hurting quality. What's your plan?

30-second answerSay your answer out loud first, then reveal.

Plan

WeekActionTypical impact
1–2Cost attribution + per-request distributions; identify the top 3 features (often ~80% of spend)Visibility
2–3Remove waste: retry storms, agent loops, duplicated calls, abuse; cap max_tokens5–20%
3–5Prompt caching (restructure prompts so the static prefix comes first); semantic/exact caches for repeated queries10–30% on input-heavy features
4–7Context diet: fewer retrieved chunks with a reranker, summarised history, shorter system prompts10–25%
5–9Model routing: small model for classification, extraction, simple Q&A; cascade to a large model on low confidence20–50% on routed traffic
6–10Move offline jobs to batch APIs~50% on those jobs
8–12Distil or fine-tune a small model for the highest-volume narrow task; consider self-hosting if volume justifies itLarge for that task

Guardrails: eval gates for every change (quality must stay within agreed tolerance), canary rollouts, a cost-per-outcome dashboard, and a weekly progress report.

Interview signal. A quantified, measured, quality-protected plan, rather than "switch to a cheaper model".

Little by little, you're building something great.