Benchmarks vs custom evaluation
A benchmark is a fixed public test set with a fixed scoring rule that is run the same way for every model, while a custom evaluation is a test you build for your own application from your own data.
Last updated: 09 Oct, 2026
LLM evaluation compared an LLM to a job candidate with past marks and an interview. The past marks are benchmarks. The interview is custom evaluation. They answer different questions, and an LLM project needs both, at different moments.
Choosing a model, then testing the application
This part of the video starts at 1:26:44. It walks through a project in the order the work happens.
- Choose the model. The first question is which LLM to use. A voice app needs low latency, because the conversation happens in real time. An app that reads images needs a multimodal model. At this step you compare models as they are, before your application exists.
- Build the application. Prompts, retrieval, tools: the custom architecture for your use case.
- Test the application. After some good-looking responses comes the need to test it properly. Ordinary tests check code. The quality of an LLM response needs an evaluation built for your case, which the video calls custom evaluation.
The slide in the clip puts the two side by side as "Two things you can evaluate". Evaluating the model is a research step: the benchmarks and leaderboards already exist, and you read them. Evaluating the application is where a developer's evaluation work goes, because nobody else can do it for your data.
What each benchmark measures
A benchmark is a set of questions with known answers and one agreed way of scoring. Because every model sits the same test, the scores can be compared. These are the ones on the slide:
| Benchmark | Area | What it holds |
|---|---|---|
| MMLU | Knowledge | Multiple-choice questions across 57 subjects, from elementary mathematics to law. |
| GSM8K | Maths | 8.5K grade-school maths word problems that need several steps. |
| MATH | Maths | 12,500 competition mathematics problems with step-by-step solutions. |
| HumanEval | Code | 164 Python functions to write from a docstring, checked by unit tests. |
| MBPP | Code | 974 basic Python programming tasks an entry-level programmer can solve. |
| HellaSwag | Commonsense | Picking the most likely next sentence after a short description of an event. |
| ARC | Science | The AI2 Reasoning Challenge: 7,787 grade-school science questions. The name is shared with ARC-AGI, a set of abstract grid puzzles. |
| GPQA | Science | 448 multiple-choice questions in biology, physics and chemistry, written by experts to be hard to answer with a web search. |
| TruthfulQA | Truthfulness | 817 questions that tempt a model into repeating a common misconception. |
| IFEval | Instructions | About 500 prompts with instructions that code can verify, such as "write in more than 400 words". |
A leaderboard collects such results in one ranking. LMArena ranks models from people's votes between two anonymous answers to the same prompt. HELM from Stanford runs models over many scenarios and reports several metrics for each. Artificial Analysis compares hosted models on quality, speed and price.
Reference-based metrics
The slide's third group scores a model's text against a reference text written by a person. No LLM judge is involved, only a formula:
- BLEU. How many word sequences of the output also appear in the reference (precision). It comes from machine translation.
- ROUGE. How many word sequences of the reference appear in the output (recall). It comes from summarization.
- METEOR. Word overlap that also counts stems and synonyms as matches.
- Exact match and token F1. Whether the answer is the reference, or how many of its words are. Common in question answering.
- BERTScore and BLEURT. A trained model compares the meaning of the two texts instead of their words.
- Perplexity. How surprised a model is by a text: the exponential of the average negative log-likelihood per token. Lower is better. It needs the model's token probabilities, so it is used while training or fine-tuning a model.
These metrics are cheap and repeatable. Their weakness is the reference: two answers can both be right and share few words, and word overlap then scores a correct answer low.
What a custom evaluation needs
The right panel of the slide starts from one sentence: no public benchmark fits your data. MMLU does not know TechNest's return policy. So you build three things:
- A golden dataset. Your own labelled examples: the questions your users ask, with the answers an expert expects. Goldens builds one.
- Task-specific metrics. What "good" means for your use case, written down as scoring rules. RAGAS metrics introduces five for a RAG app.
- An LLM as a judge. A model that reads each output and applies the metric, because answers in free text cannot be checked with
==. LLM as a judge sets one up.
The slide closes with four names. Two are frameworks that supply metrics and the code to run them: RAGAS and DeepEval. G-Eval is one metric, in which an LLM scores an output against criteria you write. The RAG triad is a set of three checks for a RAG app, drawn in RAGAS metrics. The video uses the RAGAS metrics.
Benchmarks vs custom evaluation
| Benchmark | Custom evaluation | |
|---|---|---|
| Question it answers | Which model is better in general? | Is my application doing its job? |
| Test data | Public, the same for everyone | Yours: your documents, your users' questions |
| Who builds it | Researchers, once | Your team, and it keeps growing |
| Scoring | A fixed rule, often multiple choice or unit tests | Metrics you choose, often scored by an LLM judge |
| When you use it | While picking a model | Before every release, and after every bad answer |
| What it cannot tell you | How the model behaves on your data | How the model ranks against others in general |
Where you use benchmarks and custom evaluation
- Shortlisting models. Benchmarks and leaderboards narrow twenty models down to two or three worth trying.
- Deciding between the shortlisted models. Run each one inside your application on your golden dataset. The model with the higher benchmark score does not always win here.
- Guarding a live application. Only custom evaluation notices that last week's prompt change broke the answers about refunds.
Related
- Previous: LLM evaluation
- Next: Goldens
- Reference: RAGAS metrics overview
- Open the model card or release post of the LLM you use and list which of the ten benchmarks in the table it reports. Note which areas it reports nothing for.
- Write down three questions about your own product or documents that no public benchmark could contain. They are the first rows of your golden dataset.
- Take one answer your application gave and a reference answer you write yourself, and count the words they share. Then decide whether the answer was correct. The two results often disagree, which is the weakness of word-overlap metrics.
Every expert started right here.