1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your offline eval scores improved after a change, but user complaints increased. What's going on?
30-second answerSay your answer out loud first, then reveal.
Investigation
- Read the complaints and traces: what exactly is wrong? Group them with error analysis (Q6).
- Coverage check: are those query types, languages or user segments in the eval set? Compare topic distributions (embedding clusters) between production and eval data.
- Metric mismatch: maybe the change made answers longer and more cautious. The judge rewarded thoroughness, but mobile users find it annoying. Or more refusals (safe but unhelpful).
- Overfitting: many prompt iterations against the same set can produce gains that don't generalise. Check against a held-out set.
- Judge validity: re-validate the judge against fresh human labels.
- Non-quality factors: latency increased, or the UI changed at the same time.
Fixes
- Add complaint cases and new segments to the dataset.
- Add metrics for the missing dimensions (conciseness, refusal rate, latency).
- Keep a held-out set and rotate it.
- Use online A/B tests with user-centric metrics before full rollout.
Lesson: evals are a model of user satisfaction, and models need validation too.
Related
Little by little, you're building something great.