DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

LLM evaluation

LLM evaluation is the process of checking an LLM app's answers against what they should be, using test questions with known answers and scoring rules, called metrics, that turn each answer into a number.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

With the key working, the next question is what to check. The video explains evaluation with an analogy everyone has lived through, and every DeepEval name in this course maps onto one part of it.

Evaluation as an examination · from the Production RAG Live Marathon · 3:00:48 to 3:04:53

Marking an app like an exam

The video's analogy is a basic examination. Evaluation is everywhere: a board marks your exam, an interviewer evaluates you for a job. An LLM app works on your behalf like an employee, so you want its results to be up to the mark too.

In an exam the student reads the questions and writes the answers. The teacher, who has the expertise, takes the paper, compares each answer with the correct one and gives marks. In an evaluation pipeline the student is the RAG app: it writes the answers, the real data. The teacher knows the reference answer, compares, and gives the result.

In the examIn LLM evaluationDeepEval name
The questionWhat the user asksinput
The student's answerWhat the app answeredactual_output
The teacher's correct answerThe expected answer, the ground truthexpected_output
The marking schemeA scoring rulea metric
The teacherAn LLM that applies the rulethe judge, model=
The pass markThe lowest score that passesthreshold

Comparing answers as strings

Before reaching for a judge, try the obvious rules. Here are three answers to "What is TechNest's return window?": two correct, one wrong. The first rule asks whether the answer equals the expected one; the second measures how many of the expected words it uses.

Example
expected = "Returns are accepted within 30 days of purchase."
answers = {
    "correct, same words": "Returns are accepted within 30 days of purchase.",
    "correct, own words": "You have 30 days from the day you bought it to send it back.",
    "wrong number": "Returns are accepted within 60 days of purchase.",
}


def overlap(answer, expected):
    """The share of the expected answer's words that the answer also uses."""
    said, wanted = set(answer.lower().split()), set(expected.lower().split())
    return len(said & wanted) / len(wanted)


for label, answer in answers.items():
    print(f"{label:20} equal={answer == expected!s:5}  overlap={overlap(answer, expected):.2f}")

What the two rules got wrong

  • Equality passes only the answer that copies the expected words. The correct answer in other words fails it.
  • Word overlap gives the correct paraphrase 0.25 and the wrong answer 0.88, because the wrong answer changes one word, the one that matters.
  • Both rules read wording, not meaning. Telling a right answer from a wrong one needs something that reads like the teacher does, which is why most DeepEval metrics ask an LLM to judge.

Evaluating a model vs evaluating your app

Evaluating a modelEvaluating your app
Question askedHow capable is this LLM in general?Does my app answer my users correctly?
Test questionsPublic benchmarks such as MMLUYour own goldens, written for your domain
Who builds itModel labs and leaderboardsYou
In DeepEvaldeepeval.benchmarksTest cases and metrics, the rest of this course

When to run an evaluation

  • Before a release, so a wrong answer shows up as a low score instead of a support ticket.
  • After any change to the prompt, the model or the retrieval, to see whether it helped or hurt.
  • In CI on every pull request, with a small set of questions that must always pass.
Watch out. Write the expected answer from the source, not from what the app said. An expected answer copied from the app's own output only records what the app already does, so every score comes out high and the evaluation catches nothing.
Try it yourself
  • Add a fourth answer, "Returns are accepted within 30 days of purchase!", and see which rule the exclamation mark breaks.
  • Change overlap to compare lower-cased words with punctuation removed, and check whether the wrong answer still scores higher than the correct paraphrase.

You understood something today that you didn't yesterday.