Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q28IntermediateConcept

How do you reduce an agent's cost?

30-second answerSay your answer out loud first, then reveal.

Why agents are expensive: each step re-sends the entire growing context. A 20-step run with a 30K-token context processes around 600K input tokens, before counting sub-agents.

Levers

LeverTypical impactNotes
Prompt cachingLarge reduction on cached input tokensPut static content first; keep the prefix stable (no timestamps at the top!)
Model routing / cascadesBigCheap model first; escalate to strong model on low confidence
Context trimming / compactionMedium–bigQ17
Fewer steps (better tools)Medium–bigAlso improves latency and reliability
Response/semantic cachingDepends on repetitionCareful with personalised or stale answers
Batch API for offline jobsOften ~50% cheaperFor non-interactive workloads
Output length limitsMediumOutput tokens are priced higher than input
Budgets / kill switchesPrevents disastersPer-run and per-user limits

Metric: cost per successful task. A cheaper model that fails twice as often and needs retries or human escalation can cost more overall.

Follow-ups to expect

  • How does prompt caching work? The provider stores the processed prefix (KV cache). Later requests with an identical prefix are cheaper and faster. Any change early in the prompt invalidates everything after it.

Slow is fine. Stopping is the only problem.