1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you reduce an agent's cost?
30-second answerSay your answer out loud first, then reveal.
Why agents are expensive: each step re-sends the entire growing context. A 20-step run with a 30K-token context processes around 600K input tokens, before counting sub-agents.
Levers
| Lever | Typical impact | Notes |
|---|---|---|
| Prompt caching | Large reduction on cached input tokens | Put static content first; keep the prefix stable (no timestamps at the top!) |
| Model routing / cascades | Big | Cheap model first; escalate to strong model on low confidence |
| Context trimming / compaction | Medium–big | Q17 |
| Fewer steps (better tools) | Medium–big | Also improves latency and reliability |
| Response/semantic caching | Depends on repetition | Careful with personalised or stale answers |
| Batch API for offline jobs | Often ~50% cheaper | For non-interactive workloads |
| Output length limits | Medium | Output tokens are priced higher than input |
| Budgets / kill switches | Prevents disasters | Per-run and per-user limits |
Metric: cost per successful task. A cheaper model that fails twice as often and needs retries or human escalation can cost more overall.
Follow-ups to expect
- How does prompt caching work? The provider stores the processed prefix (KV cache). Later requests with an identical prefix are cheaper and faster. Any change early in the prompt invalidates everything after it.
Related
Slow is fine. Stopping is the only problem.