1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How are LLMs evaluated? Explain perplexity, common benchmarks, and data contamination.
30-second answerSay your answer out loud first, then reveal.
Perplexity
PPL = exp( −(1/N) Σ log p(token_i | previous tokens) )- Lower is better. PPL = 10 means the model is, on average, as uncertain as choosing uniformly among 10 tokens.
- Only comparable between models with the same tokenizer on the same text.
Benchmark families (examples)
| Category | Examples |
|---|---|
| Knowledge / multitask | MMLU, MMLU-Pro, GPQA |
| Maths | GSM8K, MATH, AIME problems |
| Code | HumanEval, MBPP, LiveCodeBench, SWE-bench |
| Instruction following | IFEval |
| Chat preference | Chatbot Arena (human votes), LLM-judge suites |
| Long context | Needle-in-a-haystack, RULER |
| Agents / tools | τ-bench, terminal and web tasks |
Problems with benchmarks
- Contamination: test questions appear in web-scraped training data. Mitigations: held-out, fresh or dynamic benchmarks (e.g. problems published after the training cutoff), canary strings, n-gram overlap checks.
- Saturation: top models all score above 90%, so the benchmark no longer separates them.
- Mismatch: benchmark ≠ your task, prompt format sensitivity, LLM-judge biases.
Practical answer. "Benchmarks shortlist models; my own eval set on real tasks makes the decision."
Related
Slow is fine. Stopping is the only problem.