Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q23IntermediateConcept

How do you A/B test LLM features in production?

30-second answerSay your answer out loud first, then reveal.

Steps

  1. Offline gate first: only variants that pass offline evals go to A/B.
  2. Hypothesis + metrics: e.g. "new retrieval increases resolution rate by 3% without raising escalations or cost per conversation > 10%."
  3. Randomisation unit: user/account level; sticky assignment.
  4. Sample size / power: compute from baseline rate and minimum detectable effect. LLM features often have high variance, so plan for enough traffic.
  5. Guardrails: latency p95, error rate, safety flags, cost. Auto-stop if breached.
  6. Duration: at least 1–2 weeks for weekly seasonality; watch novelty effects.
  7. Analysis: segment by user type, language and query type, since averages hide regressions.

LLM-specific twists

  • Online quality measurement: sample conversations from both arms for blind LLM-judge or human grading.
  • Interleaving for ranking or search changes (more sensitive than A/B).
  • Cost as a first-class metric: a variant can win on quality but lose on unit economics.
  • Model provider changes can affect both arms. Pin versions during the experiment.

Common mistakes

  • Per-request randomisation in chat (the user sees different personalities between turns).
  • Declaring victory from a 2-day test on thumbs-up rate.

Slow is fine. Stopping is the only problem.