Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q4EasyConcept

What are BLEU, ROUGE, BERTScore, exact match and F1? Why are they often insufficient for LLM apps?

30-second answerSay your answer out loud first, then reveal.
MetricMeasuresGood forWeakness
Exact matchIdentical (normalised) answerShort factual answers, labelsZero credit for paraphrases
Token F1Word overlap precision/recallExtractive QASurface-level
BLEUn-gram precision vs referencesMachine translation (historical)Poor for open-ended tasks
ROUGE-LLongest common subsequence / n-gram recallSummarisation (historical)Rewards copying, not correctness
BERTScoreEmbedding similarityParaphrase-tolerant comparisonCan't detect factual errors with similar wording
Example failure. Reference "The refund window is 30 days." Output: "The refund window is 13 days." That gets high ROUGE and BERTScore despite being wrong.

When they're still useful: quick regression signals, translation benchmarks, extraction with exact values, and as cheap features alongside other checks.

Interview line. "For LLM apps I prefer task-specific criteria: exact checks where possible, plus calibrated rubric judging for meaning and faithfulness."

This is what real progress feels like.