Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q31IntermediateScenario

Your agent worked great in the demo, but fails about 30% of the time in production. How do you approach it?

30-second answerSay your answer out loud first, then reveal.

Process

  1. Define failure precisely. User thumbs-down? Escalations? Wrong actions? Timeouts? You need a measurable signal, and different signals point to different problems.
  2. Collect traces for failed and successful runs (Q23).
  3. Error analysis. Read 50–100 failed traces. Write open notes, then group them into categories, e.g.:
    • 35% wrong or insufficient retrieval
    • 20% tool errors not handled (timeouts, auth)
    • 15% ambiguous user requests where it guessed instead of asking
    • 15% wrong tool chosen
    • 10% hit max steps
    • 5% genuine model reasoning errors
  4. Fix by category, biggest and cheapest first:
    • Retrieval: chunking, hybrid search, reranking, metadata filters.
    • Tool errors: retries and better error messages.
    • Ambiguity: a clarification policy.
    • Tool confusion: descriptions and consolidation (Q19).
  5. Build the eval set from these real failures, and measure each fix against it before shipping.
  6. Ship safely: shadow mode, canary, feature flags, monitoring (Q43).
  7. Close the loop: continuously sample new production failures into the dataset.

Classic causes of demo-to-prod gaps: distribution shift in inputs, real data larger than test data (context overflow), rate limits, concurrency, auth per user, and multilingual or messy text.

Common mistakes

  • Immediately swapping to a bigger model or rewriting prompts without knowing the failure categories.

Little by little, you're building something great.