Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q28IntermediateConcept

How are LLMs evaluated? Explain perplexity, common benchmarks, and data contamination.

30-second answerSay your answer out loud first, then reveal.

Perplexity

text
PPL = exp( −(1/N) Σ log p(token_i | previous tokens) )
  • Lower is better. PPL = 10 means the model is, on average, as uncertain as choosing uniformly among 10 tokens.
  • Only comparable between models with the same tokenizer on the same text.

Benchmark families (examples)

CategoryExamples
Knowledge / multitaskMMLU, MMLU-Pro, GPQA
MathsGSM8K, MATH, AIME problems
CodeHumanEval, MBPP, LiveCodeBench, SWE-bench
Instruction followingIFEval
Chat preferenceChatbot Arena (human votes), LLM-judge suites
Long contextNeedle-in-a-haystack, RULER
Agents / toolsτ-bench, terminal and web tasks

Problems with benchmarks

  • Contamination: test questions appear in web-scraped training data. Mitigations: held-out, fresh or dynamic benchmarks (e.g. problems published after the training cutoff), canary strings, n-gram overlap checks.
  • Saturation: top models all score above 90%, so the benchmark no longer separates them.
  • Mismatch: benchmark ≠ your task, prompt format sensitivity, LLM-judge biases.
Practical answer. "Benchmarks shortlist models; my own eval set on real tasks makes the decision."

Slow is fine. Stopping is the only problem.