Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q42HardConcept

If each step of an agent is 95% reliable, what happens on a 20-step task? How do you build reliable agents anyway?

30-second answerSay your answer out loud first, then reveal.

The math

Per-step success10 steps20 steps50 steps
95%60%36%8%
99%90%82%61%
99.9%99%98%95%

Strategies

  1. Fewer steps: more powerful, higher-level tools (one create_invoice_from_po instead of eight API calls). Hard-code deterministic parts as workflow.
  2. Higher per-step reliability: clear tool schemas, structured outputs, validation with auto-retry, good error messages, better models for critical decisions.
  3. Verification after actions: check the environment state after acting (did the file change? did the record save?). Run tests. This turns silent errors into visible ones.
  4. Recovery: let the agent see errors and retry differently. Checkpoints let you roll back to a known good state. This is why the independence assumption is pessimistic for good agents: they self-correct.
  5. Decomposition with checkpoints: split a 50-step task into 5 stages with verified outputs; a failure only reruns its stage.
  6. Human gates at high-stakes points.
  7. Measure reliability properly: pass^k (all k runs succeed) vs pass@k (any run succeeds). Production users need consistent behaviour, so pass^k matters.
Interview gold: quoting the compounding math, then explaining why verification and recovery change it.

You understood something today that you didn't yesterday.