LLM evaluation
LLM evaluation is the process of checking an LLM app's answers against what they should be, using test questions with known answers and scoring rules, called metrics, that turn each answer into a number.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
With the key working, the next question is what to check. The video explains evaluation with an analogy everyone has lived through, and every DeepEval name in this course maps onto one part of it.
Marking an app like an exam
The video's analogy is a basic examination. Evaluation is everywhere: a board marks your exam, an interviewer evaluates you for a job. An LLM app works on your behalf like an employee, so you want its results to be up to the mark too.
In an exam the student reads the questions and writes the answers. The teacher, who has the expertise, takes the paper, compares each answer with the correct one and gives marks. In an evaluation pipeline the student is the RAG app: it writes the answers, the real data. The teacher knows the reference answer, compares, and gives the result.
| In the exam | In LLM evaluation | DeepEval name |
|---|---|---|
| The question | What the user asks | input |
| The student's answer | What the app answered | actual_output |
| The teacher's correct answer | The expected answer, the ground truth | expected_output |
| The marking scheme | A scoring rule | a metric |
| The teacher | An LLM that applies the rule | the judge, model= |
| The pass mark | The lowest score that passes | threshold |
Comparing answers as strings
Before reaching for a judge, try the obvious rules. Here are three answers to "What is TechNest's return window?": two correct, one wrong. The first rule asks whether the answer equals the expected one; the second measures how many of the expected words it uses.
expected = "Returns are accepted within 30 days of purchase."
answers = {
"correct, same words": "Returns are accepted within 30 days of purchase.",
"correct, own words": "You have 30 days from the day you bought it to send it back.",
"wrong number": "Returns are accepted within 60 days of purchase.",
}
def overlap(answer, expected):
"""The share of the expected answer's words that the answer also uses."""
said, wanted = set(answer.lower().split()), set(expected.lower().split())
return len(said & wanted) / len(wanted)
for label, answer in answers.items():
print(f"{label:20} equal={answer == expected!s:5} overlap={overlap(answer, expected):.2f}")correct, same words equal=True overlap=1.00 correct, own words equal=False overlap=0.25 wrong number equal=False overlap=0.88
What the two rules got wrong
- Equality passes only the answer that copies the expected words. The correct answer in other words fails it.
- Word overlap gives the correct paraphrase 0.25 and the wrong answer 0.88, because the wrong answer changes one word, the one that matters.
- Both rules read wording, not meaning. Telling a right answer from a wrong one needs something that reads like the teacher does, which is why most DeepEval metrics ask an LLM to judge.
Evaluating a model vs evaluating your app
| Evaluating a model | Evaluating your app | |
|---|---|---|
| Question asked | How capable is this LLM in general? | Does my app answer my users correctly? |
| Test questions | Public benchmarks such as MMLU | Your own goldens, written for your domain |
| Who builds it | Model labs and leaderboards | You |
| In DeepEval | deepeval.benchmarks | Test cases and metrics, the rest of this course |
When to run an evaluation
- Before a release, so a wrong answer shows up as a low score instead of a support ticket.
- After any change to the prompt, the model or the retrieval, to see whether it helped or hurt.
- In CI on every pull request, with a small set of questions that must always pass.
Related
- Previous: Installation and setup
- Next: TechNest RAG app
- Reference: Introduction to LLM evals
- Add a fourth answer,
"Returns are accepted within 30 days of purchase!", and see which rule the exclamation mark breaks. - Change
overlapto compare lower-cased words with punctuation removed, and check whether the wrong answer still scores higher than the correct paraphrase.
You understood something today that you didn't yesterday.