AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

LLM evaluation

LLM evaluation is the practice of running an LLM application on a fixed set of test questions and scoring its outputs with defined metrics.

Last updated: 09 Oct, 2026 · OpenAI SDK 3.3

Guardrails filter requests and replies, and tracing shows what each call did, as in LLM observability with Pydantic Logfire. Neither says whether an answer was right. Evaluation does: it turns "the answers look fine" into numbers you can compare before and after every change.

A raw RAG app vs an app with evaluation

The TechNest evaluation app · from the Complete AI Security Course in 8 Hours video · 1:18:07 to 1:22:42

This part of the video starts at 1:18:07. The app in the clip answers with llama-3.3-70b-versatile and judges with llama-3.1-8b-instant, both since retired on Groq; the code below answers with qwen/qwen3.8-27b.

The evaluation module of the video opens with a document chat: a PDF goes in, a question gets an answer, and the source chunks are listed under it. A RAG app (retrieval-augmented generation) fetches the passages closest to a question and lets an LLM write the answer from them. The video calls this first app a raw RAG: nothing on the screen says whether the answer is good, and the only check is to read it yourself.

The second app in the clip is the one this part is built around. It belongs to TechNest, a made-up online electronics store, and has four tabs:

  • Catalog. The knowledge base: 15 entries (products, policies and FAQs) that the RAG system retrieves from. Each entry has an ID, a category, a title and the content.
  • Goldens. The questions with the answers that should come back. They are "the basis on which you are going to evaluate your application".
  • Run Evaluation. First the RAG pipeline answers every golden question. Then each answer is scored with metrics from the RAGAS framework, by an LLM used as a judge.
  • Results. One score per question and per metric, plus the averages.

For the question "What is TechNest's return policy?" the app shows three things side by side: the chunks it retrieved, the answer it wrote, and the reference answer, "the answer which should come". With those in hand it can ask the questions a reader of the raw app could only guess at: is the answer faithful to the chunks, is it relevant to the question, are the retrieved chunks the right ones?

The video gives the reason this matters once an app is live: "this information will help you figure out where is my application lagging". When users find a question the chatbot answers badly, that question is added as a new test case, and from then on every version of the app has to pass it.

Evaluation as an interview

The candidate analogy for evaluation · from the Complete AI Security Course in 8 Hours video · 1:23:27 to 1:26:35

This part of the video starts at 1:23:27. A benchmark is a fixed public test set that is scored the same way for every model; the model release named in the clip is the example of that week.

The video explains evaluation with a job candidate. No company hires without an interview. A candidate walks in with two kinds of information:

  • Past marks. The board lists 10th: 82%, 12th: 85% and CGPA: 9. These are predefined scores, already there before the company meets the candidate. They are useful, and they are not enough to hire on.
  • The interview. A test with a score, a reasoning ability test and a one-on-one interview. Here the company judges the candidate on the use case it wants to put them in, for example a Gen AI engineering role.

An LLM comes with the same two kinds of information. Its past marks are benchmarks, the public scores published with every new model. Its interview is the evaluation you run for your own use case, where a person or another LLM reads the outputs and judges them: "like a human is judging a human", here "an LLM is judging an LLM". And because an LLM judge can be wrong too, a person stays in the review loop and reads a share of what the judge decided.

Two rows that mirror each other. A candidate has past marks (10th 82%, 12th 85%, CGPA 9) and an interview for the role (a test, reasoning ability, a one-on-one interview). An LLM has benchmarks, fixed public test sets run the same way for every model, and custom evaluation on your own questions, data and metrics, scored by a human or an LLM judge.

Benchmarks vs custom evaluation takes the two columns apart. The rest of this page builds the thing to be evaluated.

The TechNest app under test

Evaluation needs an app to evaluate. The video's repository holds the TechNest support bot: a catalog, a retriever and a generator. The same three steps fit in one short file.

The TechNest app as a flow: the question is embedded with gemini-embedding-2, the retriever picks the top 3 catalog entries by similarity from a catalog of 6 entries embedded once, and the generator writes the answer from those chunks. The retrieved contexts and the response are the two outputs that evaluation reads.

The video's repository wraps these steps in a Retriever class and a Generator class, embeds through LangChain's Gemini wrapper and answers with llama-3.3-70b-versatile. The code here keeps the same steps and the same system prompt as three functions, calls gemini-embedding-2 through google-genai with one text per call, and answers with qwen/qwen3.8-27b, because the video's model is retired. One line is added to the prompt to keep answers to a few sentences, and the catalog holds 6 of the repository's 15 entries (three products and three policies) so that a run needs few embedding calls. The support e-mail address in the warranty entry is replaced by the words "TechNest support".

Embedding a text with Gemini

An embedding turns a text into a list of numbers, so that texts with similar meaning get similar lists. embed sends one text to Gemini and scales the result to length 1. Each catalog entry is embedded once, as its title plus its content, the way the video's retriever does it.

python
gemini = genai.Client(api_key=os.environ["GEMINI_API_KEY"])


def embed(text):
    values = gemini.models.embed_content(model=EMBED_MODEL, contents=text).embeddings[0].values
    vector = np.array(values)
    return vector / np.linalg.norm(vector)           # length 1, so a dot product is the cosine


DOC_VECTORS = np.array([embed(f"{title}. {content}") for title, content in CATALOG])

Retrieving the top three chunks

@ multiplies the question's vector with every catalog vector at once, which gives one cosine similarity per entry. The three highest are returned, best first.

python
def retrieve(question, top_k=3):
    scores = DOC_VECTORS @ embed(question)            # one similarity score per catalog entry
    best = np.argsort(scores)[::-1][:top_k]           # highest first
    return [CATALOG[i][1] for i in best]

Generating the answer from the chunks

The numbered chunks and the question go in the user message. The system prompt, the repository's own, tells the model to answer from the context only and to say so when the context is not enough.

python
def generate(question, contexts):
    context_block = "\n\n".join(f"[{i + 1}] {c}" for i, c in enumerate(contexts))
    messages = [{"role": "system", "content": SYSTEM_PROMPT},
                {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"}]
    reply = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0, max_tokens=300)
    return reply.choices[0].message.content.strip()

Asking the app the video's question

The whole file, with the question from the clip at the end. It needs GROQ_API_KEY and GEMINI_API_KEY, set up in Installing Python for AI security.

ExampleAPI keyFrom the video, run on Groq
import os
import textwrap

import numpy as np
from google import genai
from openai import OpenAI

CATALOG = [
    ("ProBook X1 Laptop", "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."),
    ("PixelPhone 15", "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."),
    ("SoundPods Pro", "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."),
    ("Return Policy", "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."),
    ("Shipping Policy", "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."),
    ("Warranty Policy", "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact TechNest support with your order number and a description of the issue."),
]
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Answer in at most three plain sentences, with no lists and no bold text."""
CHAT_MODEL = "qwen/qwen3.8-27b"
EMBED_MODEL = "gemini-embedding-2"

groq = OpenAI(base_url="https://api.groq.com/openai/v1", api_key=os.environ["GROQ_API_KEY"])
gemini = genai.Client(api_key=os.environ["GEMINI_API_KEY"])


def embed(text):
    values = gemini.models.embed_content(model=EMBED_MODEL, contents=text).embeddings[0].values
    vector = np.array(values)
    return vector / np.linalg.norm(vector)           # length 1, so a dot product is the cosine


DOC_VECTORS = np.array([embed(f"{title}. {content}") for title, content in CATALOG])


def retrieve(question, top_k=3):
    scores = DOC_VECTORS @ embed(question)            # one similarity score per catalog entry
    best = np.argsort(scores)[::-1][:top_k]           # highest first
    return [CATALOG[i][1] for i in best]


def generate(question, contexts):
    context_block = "\n\n".join(f"[{i + 1}] {c}" for i, c in enumerate(contexts))
    messages = [{"role": "system", "content": SYSTEM_PROMPT},
                {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"}]
    reply = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0, max_tokens=300)
    return reply.choices[0].message.content.strip()


def answer(question):
    contexts = retrieve(question)
    return generate(question, contexts), contexts


question = "What is TechNest's return policy?"
response, contexts = answer(question)
for number, chunk in enumerate(contexts, 1):
    print(f"chunk {number}: {chunk[:62]}...")
print()
print(textwrap.fill(response, 100))

What the app returned

  • The retriever put the return policy first. Chunk 1 is the return policy entry, the one the question is about. Chunks 2 and 3 are the warranty and shipping policies: the closest of the remaining entries, and of no use for this question.
  • The answer follows chunk 1. The 30 days, the original packaging, who pays for return shipping, the 5 to 7 business days and the non-refundable items are all in the return policy entry.
  • You know that only because you compared them. Nothing in the output marks the answer as checked. You read the chunk, read the answer and matched them sentence by sentence.

Why an answer that looks right is not a test

The answer above reads well. To call it correct you would still have to check several things by hand:

  • Is every sentence taken from the chunks? A fluent sentence the chunks do not contain reads exactly like one they do.
  • Is anything the policy says missing? You only notice a gap if you know the full answer.
  • Were the right chunks retrieved? Two of the three chunks are about other policies. Here they did no harm; for another question they might.
  • Will it still be right tomorrow? A new prompt, a new model or a bigger catalog can change any of the above, for this question or for one you did not try.

Evaluation answers each of these with a number, for a whole set of questions at once. Goldens builds the question set, LLM as a judge sets up the reader that does the checking, and RAGAS metrics names the checks.

Reading answers by eye vs evaluation

Reading answers by eyeEvaluation
What is checkedThe few questions you happened to tryA fixed set of questions, the same every run
Who decidesYour impression at the timeWritten metrics, scored the same way each run
After a changeStart again from memoryRun the set again and compare the numbers
What it findsObvious failuresWhich step is weak: retrieval or generation
CostYour time, every timeModel calls, plus the time to write the questions once

Where you use LLM evaluation

  • Before a release. A prompt, model or retriever change ships only if the scores on the question set did not drop.
  • After a bad answer in production. The question becomes a test case, so the same failure cannot come back unnoticed.
  • When choosing between two designs. Top 3 or top 5 chunks, one model or another: run both on the same questions and compare.
Watch out. A demo that answers three hand-picked questions well has shown three answers, not the app. The failures that matter are in the questions nobody tried, which is why the question set is written down first and run every time.
Try it yourself
  • Change top_k=3 to top_k=1 in retrieve and run again: only the return policy chunk is printed, and the answer is written from it alone.
  • Ask "What is the price of the PixelPhone 15?" instead: the first chunk becomes the PixelPhone 15 entry and the answer gives its price, $899.
  • Ask "Do you ship to Canada?". None of the six entries covers shipping outside the continental US. Read the answer and decide for yourself whether the app handled the missing information well; that judgement is what the next lessons turn into a score.

This is what real progress feels like.