1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Pointwise vs pairwise evaluation: when do you use each? How do arena-style rankings work?
30-second answerSay your answer out loud first, then reveal.
| Pointwise | Pairwise | |
|---|---|---|
| Question | "Is this good enough?" | "Is A or B better?" |
| Output | Absolute score / pass rate | Win rate / preference |
| Sensitivity to small differences | Lower | Higher |
| Use cases | Release gates, monitoring, SLAs | Prompt/model A vs B, preference data for DPO |
| Biases | Leniency, scale drift | Position bias, length bias |
Pairwise best practices
- Swap order and run both orders; count ties when the verdicts conflict.
- Control for length: instruct the judge, or compare length-normalised outputs.
- Report win / tie / loss rates with confidence intervals.
Elo and Bradley–Terry
Elo / Bradley–Terry (intuition): each system has a latent strength. The probability A beats B depends on the strength difference. Fitting this to many pairwise outcomes gives a ranking across many systems, as used by Chatbot Arena–style leaderboards.
Practical combo: pairwise to choose between candidate prompts or models; pointwise on the winner to confirm it meets absolute quality bars.
Related
Slow is fine. Stopping is the only problem.