Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q46HardScenario

A model ranks top on public benchmarks but performs poorly on your company's task. Explain why, and how you'd choose a model properly.

30-second answerSay your answer out loud first, then reveal.

Why the mismatch

  1. Contamination: benchmark items appear in training data, so scores reflect memorisation.
  2. Distribution shift: your inputs (Hinglish chats, messy invoices, internal jargon, long documents) differ from benchmark questions.
  3. Task shape: benchmarks are short-answer; your task is multi-turn, tool-using, long-context, or needs strict formats.
  4. Prompt sensitivity: a model may need different prompting styles; a single shared prompt isn't fair.
  5. Benchmark-targeted tuning: models optimised for popular leaderboards.
  6. Metric mismatch: accuracy on multiple choice vs your need for calibrated abstention, latency or cost.

Proper model selection process

  1. Define success: task metrics (exact match, F1, rubric scores), plus latency (TTFT, total), cost per task, safety constraints.
  2. Build the eval set: 100–500 real examples covering common cases, edge cases and adversarial inputs, with references or rubrics.
  3. Fair testing: tune prompts per model (briefly), use the same retrieval/context, run 3+ trials for variance.
  4. Scoring: automatic where possible; calibrated LLM judges or humans for open-ended outputs.
  5. Analyse failures, not just averages: some models fail catastrophically on certain segments.
  6. Total cost of ownership: API vs self-hosted, rate limits, data policies.
  7. Re-run on every new model release, since your eval is a reusable asset.

Little by little, you're building something great.