AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Benchmarks vs custom evaluation

A benchmark is a fixed public test set with a fixed scoring rule that is run the same way for every model, while a custom evaluation is a test you build for your own application from your own data.

Last updated: 09 Oct, 2026

LLM evaluation compared an LLM to a job candidate with past marks and an interview. The past marks are benchmarks. The interview is custom evaluation. They answer different questions, and an LLM project needs both, at different moments.

Choosing a model, then testing the application

Evaluating the model vs evaluating the application · from the Complete AI Security Course in 8 Hours video · 1:26:44 to 1:30:08

This part of the video starts at 1:26:44. It walks through a project in the order the work happens.

  1. Choose the model. The first question is which LLM to use. A voice app needs low latency, because the conversation happens in real time. An app that reads images needs a multimodal model. At this step you compare models as they are, before your application exists.
  2. Build the application. Prompts, retrieval, tools: the custom architecture for your use case.
  3. Test the application. After some good-looking responses comes the need to test it properly. Ordinary tests check code. The quality of an LLM response needs an evaluation built for your case, which the video calls custom evaluation.

The slide in the clip puts the two side by side as "Two things you can evaluate". Evaluating the model is a research step: the benchmarks and leaderboards already exist, and you read them. Evaluating the application is where a developer's evaluation work goes, because nobody else can do it for your data.

Two panels side by side. Evaluating the model asks which model is better and uses ready-made material: benchmarks (MMLU for knowledge, GSM8K and MATH for maths, HumanEval and MBPP for code, HellaSwag for commonsense, ARC and GPQA for science, TruthfulQA for truthfulness, IFEval for instructions), leaderboards (LMArena, HELM, Artificial Analysis) and reference-based metrics (BLEU, ROUGE, METEOR, exact match, BERTScore, BLEURT, perplexity). Evaluating the application asks whether your RAG app or agent does its job: no public benchmark fits your data, so you build a golden dataset, task-specific metrics and an LLM judge, with frameworks such as RAGAS and DeepEval.

What each benchmark measures

A benchmark is a set of questions with known answers and one agreed way of scoring. Because every model sits the same test, the scores can be compared. These are the ones on the slide:

BenchmarkAreaWhat it holds
MMLUKnowledgeMultiple-choice questions across 57 subjects, from elementary mathematics to law.
GSM8KMaths8.5K grade-school maths word problems that need several steps.
MATHMaths12,500 competition mathematics problems with step-by-step solutions.
HumanEvalCode164 Python functions to write from a docstring, checked by unit tests.
MBPPCode974 basic Python programming tasks an entry-level programmer can solve.
HellaSwagCommonsensePicking the most likely next sentence after a short description of an event.
ARCScienceThe AI2 Reasoning Challenge: 7,787 grade-school science questions. The name is shared with ARC-AGI, a set of abstract grid puzzles.
GPQAScience448 multiple-choice questions in biology, physics and chemistry, written by experts to be hard to answer with a web search.
TruthfulQATruthfulness817 questions that tempt a model into repeating a common misconception.
IFEvalInstructionsAbout 500 prompts with instructions that code can verify, such as "write in more than 400 words".

A leaderboard collects such results in one ranking. LMArena ranks models from people's votes between two anonymous answers to the same prompt. HELM from Stanford runs models over many scenarios and reports several metrics for each. Artificial Analysis compares hosted models on quality, speed and price.

Reference-based metrics

The slide's third group scores a model's text against a reference text written by a person. No LLM judge is involved, only a formula:

  • BLEU. How many word sequences of the output also appear in the reference (precision). It comes from machine translation.
  • ROUGE. How many word sequences of the reference appear in the output (recall). It comes from summarization.
  • METEOR. Word overlap that also counts stems and synonyms as matches.
  • Exact match and token F1. Whether the answer is the reference, or how many of its words are. Common in question answering.
  • BERTScore and BLEURT. A trained model compares the meaning of the two texts instead of their words.
  • Perplexity. How surprised a model is by a text: the exponential of the average negative log-likelihood per token. Lower is better. It needs the model's token probabilities, so it is used while training or fine-tuning a model.
Perplexity of a text of N tokens

These metrics are cheap and repeatable. Their weakness is the reference: two answers can both be right and share few words, and word overlap then scores a correct answer low.

What a custom evaluation needs

The right panel of the slide starts from one sentence: no public benchmark fits your data. MMLU does not know TechNest's return policy. So you build three things:

  • A golden dataset. Your own labelled examples: the questions your users ask, with the answers an expert expects. Goldens builds one.
  • Task-specific metrics. What "good" means for your use case, written down as scoring rules. RAGAS metrics introduces five for a RAG app.
  • An LLM as a judge. A model that reads each output and applies the metric, because answers in free text cannot be checked with ==. LLM as a judge sets one up.

The slide closes with four names. Two are frameworks that supply metrics and the code to run them: RAGAS and DeepEval. G-Eval is one metric, in which an LLM scores an output against criteria you write. The RAG triad is a set of three checks for a RAG app, drawn in RAGAS metrics. The video uses the RAGAS metrics.

Benchmarks vs custom evaluation

BenchmarkCustom evaluation
Question it answersWhich model is better in general?Is my application doing its job?
Test dataPublic, the same for everyoneYours: your documents, your users' questions
Who builds itResearchers, onceYour team, and it keeps growing
ScoringA fixed rule, often multiple choice or unit testsMetrics you choose, often scored by an LLM judge
When you use itWhile picking a modelBefore every release, and after every bad answer
What it cannot tell youHow the model behaves on your dataHow the model ranks against others in general

Where you use benchmarks and custom evaluation

  • Shortlisting models. Benchmarks and leaderboards narrow twenty models down to two or three worth trying.
  • Deciding between the shortlisted models. Run each one inside your application on your golden dataset. The model with the higher benchmark score does not always win here.
  • Guarding a live application. Only custom evaluation notices that last week's prompt change broke the answers about refunds.
Watch out. A benchmark score is a statement about a public test set. Public questions can end up in training data, and a model that tops a leaderboard has still never seen your catalog. Treat benchmarks as a filter for choosing candidates, and trust only the evaluation you ran on your own data.
Try it yourself
  • Open the model card or release post of the LLM you use and list which of the ten benchmarks in the table it reports. Note which areas it reports nothing for.
  • Write down three questions about your own product or documents that no public benchmark could contain. They are the first rows of your golden dataset.
  • Take one answer your application gave and a reference answer you write yourself, and count the words they share. Then decide whether the answer was correct. The two results often disagree, which is the weakness of word-overlap metrics.

Every expert started right here.