LLM evaluation
LLM evaluation is the systematic process of measuring how well an AI system performs against predefined objectives and expected outcomes, so you can judge its quality, effectiveness and reliability with numbers.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
A demo that answers three questions well tells you little about the thousand questions real users will ask. Evaluation replaces "it looks fine" with a score per question, per quality.
Measuring against expected outcomes
The video reads the definition slowly. Measuring and assessing mean putting numbers on how well the app performs. Predefined objectives and expected outcomes are the goldens: the questions the app must handle, each with the answer you expect. Together they tell you where the app works and where it does not.
Evaluating the LLM vs evaluating the application
An AI application has two parts: the LLM that powers it, and all the code around the LLM. The video separates the evaluation of each.
- Evaluating the LLM. Correctness, reasoning, hallucination, safety, bias and coherence of the model itself. This is mostly done for you: public benchmarks such as the leaderboards on artificialanalysis.ai and Hugging Face compare models on intelligence, speed and cost per task. You read them to choose a model.
- Evaluating the application. Whether your whole pipeline, retrieval, prompt and model together, answers your users' questions the way you expect. No benchmark knows your catalog or your policies, so you build this yourself. The video calls this the part where most of your time goes, and it is what RAGAS is for.
| LLM evaluation | Application evaluation | |
|---|---|---|
| What is tested | The model on its own | Your retrieval, prompt and model together |
| Test data | Public benchmark sets | Your own goldens |
| Who runs it | Model providers and leaderboards | You |
| Question it answers | Which model should I use? | Does my app answer my users correctly? |
Why a RAG app needs application evaluation
A strong model can still give a wrong answer inside your app. The search can bring back the wrong chunk, the prompt can let the model add details the chunk never said, or the answer can drift off the question. A benchmark score for the model cannot see any of that; a score on your own questions can.
When to evaluate your app
- Before the first release, to know the baseline quality.
- After a change to the retriever, the chunking, the prompt or the model, to see whether the change helped.
- When users report a bad answer: add that question to the goldens so it is checked from then on.
Related
- Previous: Installation and setup
- Next: TechNest RAG app
- Reference: Metrics overview
- Open artificialanalysis.ai/models and find where
gpt-oss-20b, the course's judge, sits on intelligence and speed. - Write down three questions a TechNest customer might ask that a public benchmark would never contain.
You understood something today that you didn't yesterday.