1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you A/B test LLM features in production?
30-second answerSay your answer out loud first, then reveal.
Steps
- Offline gate first: only variants that pass offline evals go to A/B.
- Hypothesis + metrics: e.g. "new retrieval increases resolution rate by 3% without raising escalations or cost per conversation > 10%."
- Randomisation unit: user/account level; sticky assignment.
- Sample size / power: compute from baseline rate and minimum detectable effect. LLM features often have high variance, so plan for enough traffic.
- Guardrails: latency p95, error rate, safety flags, cost. Auto-stop if breached.
- Duration: at least 1–2 weeks for weekly seasonality; watch novelty effects.
- Analysis: segment by user type, language and query type, since averages hide regressions.
LLM-specific twists
- Online quality measurement: sample conversations from both arms for blind LLM-judge or human grading.
- Interleaving for ranking or search changes (more sensitive than A/B).
- Cost as a first-class metric: a variant can win on quality but lose on unit economics.
- Model provider changes can affect both arms. Pin versions during the experiment.
Common mistakes
- Per-request randomisation in chat (the user sees different personalities between turns).
- Declaring victory from a 2-day test on thumbs-up rate.
Related
Slow is fine. Stopping is the only problem.