1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are BLEU, ROUGE, BERTScore, exact match and F1? Why are they often insufficient for LLM apps?
30-second answerSay your answer out loud first, then reveal.
| Metric | Measures | Good for | Weakness |
|---|---|---|---|
| Exact match | Identical (normalised) answer | Short factual answers, labels | Zero credit for paraphrases |
| Token F1 | Word overlap precision/recall | Extractive QA | Surface-level |
| BLEU | n-gram precision vs references | Machine translation (historical) | Poor for open-ended tasks |
| ROUGE-L | Longest common subsequence / n-gram recall | Summarisation (historical) | Rewards copying, not correctness |
| BERTScore | Embedding similarity | Paraphrase-tolerant comparison | Can't detect factual errors with similar wording |
Example failure. Reference "The refund window is 30 days." Output: "The refund window is 13 days." That gets high ROUGE and BERTScore despite being wrong.
When they're still useful: quick regression signals, translation benchmarks, extraction with exact values, and as cheap features alongside other checks.
Interview line. "For LLM apps I prefer task-specific criteria: exact checks where possible, plus calibrated rubric judging for meaning and faithfulness."
Related
This is what real progress feels like.