1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
A model ranks top on public benchmarks but performs poorly on your company's task. Explain why, and how you'd choose a model properly.
30-second answerSay your answer out loud first, then reveal.
Why the mismatch
- Contamination: benchmark items appear in training data, so scores reflect memorisation.
- Distribution shift: your inputs (Hinglish chats, messy invoices, internal jargon, long documents) differ from benchmark questions.
- Task shape: benchmarks are short-answer; your task is multi-turn, tool-using, long-context, or needs strict formats.
- Prompt sensitivity: a model may need different prompting styles; a single shared prompt isn't fair.
- Benchmark-targeted tuning: models optimised for popular leaderboards.
- Metric mismatch: accuracy on multiple choice vs your need for calibrated abstention, latency or cost.
Proper model selection process
- Define success: task metrics (exact match, F1, rubric scores), plus latency (TTFT, total), cost per task, safety constraints.
- Build the eval set: 100–500 real examples covering common cases, edge cases and adversarial inputs, with references or rubrics.
- Fair testing: tune prompts per model (briefly), use the same retrieval/context, run 3+ trials for variance.
- Scoring: automatic where possible; calibrated LLM judges or humans for open-ended outputs.
- Analyse failures, not just averages: some models fail catastrophically on certain segments.
- Total cost of ownership: API vs self-hosted, rate limits, data policies.
- Re-run on every new model release, since your eval is a reusable asset.
Related
Little by little, you're building something great.