Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q45HardScenario

You must choose between two LLMs for a production feature. Design the evaluation and write the decision summary.

30-second answerSay your answer out loud first, then reveal.

Evaluation design

  1. Fair prompts: a short tuning budget per model on the dev set (models differ in prompting needs).
  2. Same test set: 300–500 examples with slices; 3 runs each for variance.
  3. Metrics: task quality (judges + code checks), safety red-team pass rate, format adherence, p50/p95 latency, cost per task, refusal rate.
  4. Paired analysis: examples where they differ; McNemar test or bootstrap CIs.
  5. Operational checks: rate limits, regional availability, data retention terms, model deprecation timelines.

Decision memo (example)

Recommendation: Model B for the claims summary feature.

Evidence: Model B matches Model A on summary correctness (91.2% vs 92.0%, difference not significant, 95% CI −2.1 to +0.5 points), has fewer critical omissions (1.1% vs 1.8%), and is 2.4x faster at p95 (6.1s vs 14.7s) and 58% cheaper per summary.

Trade-offs: slightly weaker on Hindi documents (−3 points; 8% of volume), mitigated by routing Hindi documents to Model A.

Risks: newer model version, so we pin the version and run weekly regression evals; fallback to Model A configured in the gateway.

Decision needed by: Oct 20, for pilot rollout.

Every expert started right here.