Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q26IntermediateScenario

Your offline eval scores improved after a change, but user complaints increased. What's going on?

30-second answerSay your answer out loud first, then reveal.

Investigation

  1. Read the complaints and traces: what exactly is wrong? Group them with error analysis (Q6).
  2. Coverage check: are those query types, languages or user segments in the eval set? Compare topic distributions (embedding clusters) between production and eval data.
  3. Metric mismatch: maybe the change made answers longer and more cautious. The judge rewarded thoroughness, but mobile users find it annoying. Or more refusals (safe but unhelpful).
  4. Overfitting: many prompt iterations against the same set can produce gains that don't generalise. Check against a held-out set.
  5. Judge validity: re-validate the judge against fresh human labels.
  6. Non-quality factors: latency increased, or the UI changed at the same time.

Fixes

  • Add complaint cases and new segments to the dataset.
  • Add metrics for the missing dimensions (conciseness, refusal rate, latency).
  • Keep a held-out set and rotate it.
  • Use online A/B tests with user-centric metrics before full rollout.

Lesson: evals are a model of user satisfaction, and models need validation too.

Little by little, you're building something great.