1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
If each step of an agent is 95% reliable, what happens on a 20-step task? How do you build reliable agents anyway?
30-second answerSay your answer out loud first, then reveal.
The math
| Per-step success | 10 steps | 20 steps | 50 steps |
|---|---|---|---|
| 95% | 60% | 36% | 8% |
| 99% | 90% | 82% | 61% |
| 99.9% | 99% | 98% | 95% |
Strategies
- Fewer steps: more powerful, higher-level tools (one
create_invoice_from_poinstead of eight API calls). Hard-code deterministic parts as workflow. - Higher per-step reliability: clear tool schemas, structured outputs, validation with auto-retry, good error messages, better models for critical decisions.
- Verification after actions: check the environment state after acting (did the file change? did the record save?). Run tests. This turns silent errors into visible ones.
- Recovery: let the agent see errors and retry differently. Checkpoints let you roll back to a known good state. This is why the independence assumption is pessimistic for good agents: they self-correct.
- Decomposition with checkpoints: split a 50-step task into 5 stages with verified outputs; a failure only reruns its stage.
- Human gates at high-stakes points.
- Measure reliability properly: pass^k (all k runs succeed) vs pass@k (any run succeeds). Production users need consistent behaviour, so pass^k matters.
Interview gold: quoting the compounding math, then explaining why verification and recovery change it.
Related
- Previous: Q41. Your team built a 6-agent system. It costs 10x more than the old single agent, with no quality improvement. What do you do?
- Next: Q43. You need to switch the agent's model to a newer or cheaper one. After switching, some behaviours regress. How do you manage model migrations?
- LangGraph
- DeepEval
You understood something today that you didn't yesterday.