1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your agent worked great in the demo, but fails about 30% of the time in production. How do you approach it?
30-second answerSay your answer out loud first, then reveal.
Process
- Define failure precisely. User thumbs-down? Escalations? Wrong actions? Timeouts? You need a measurable signal, and different signals point to different problems.
- Collect traces for failed and successful runs (Q23).
- Error analysis. Read 50–100 failed traces. Write open notes, then group them into categories, e.g.:
• 35% wrong or insufficient retrieval
• 20% tool errors not handled (timeouts, auth)
• 15% ambiguous user requests where it guessed instead of asking
• 15% wrong tool chosen
• 10% hit max steps
• 5% genuine model reasoning errors - Fix by category, biggest and cheapest first:
• Retrieval: chunking, hybrid search, reranking, metadata filters.
• Tool errors: retries and better error messages.
• Ambiguity: a clarification policy.
• Tool confusion: descriptions and consolidation (Q19). - Build the eval set from these real failures, and measure each fix against it before shipping.
- Ship safely: shadow mode, canary, feature flags, monitoring (Q43).
- Close the loop: continuously sample new production failures into the dataset.
Classic causes of demo-to-prod gaps: distribution shift in inputs, real data larger than test data (context overflow), rate limits, concurrency, auth per user, and multilingual or messy text.
Common mistakes
- Immediately swapping to a bigger model or rewriting prompts without knowing the failure categories.
Related
Little by little, you're building something great.