RAGASragas 0.4.3 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your path

Project: TechNest support bot evaluation

The TechNest support bot evaluation is a RAGAS project that runs the RAG app over its five goldens, scores every answer on faithfulness, answer relevancy, context precision, context recall and answer correctness, and prints a report per golden and per metric.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

The overview promised the video's evaluation app, rebuilt in Python: the app, the goldens, a judge, and a score for each question on each side of quality. This is that app's two-phase pipeline in one script, followed by a deliberate failure to show the report catching it.

The two-phase pipeline

  • Phase one runs the app on every golden with the experiment from Experiments, keeping each response and its chunks.
  • Phase two scores the rows metric by metric, one sample at a time, with the video's cooldowns from Rate limits and cooldowns: 25 seconds between samples and 35 between metrics.
  • The report uses the pass bands from Evaluation results.
The TechNest evaluation suite, and the lesson each piece came from
five metrics, one at a timerowsquestionanswer + chunksascore()judgesembedswritesDatasetgoldens.csv@experimentrun_app(row)technest.answerresponse, chunksphase-one-k3.csvexperiments/Faithfulnessclaims in the chunks?Retrieval metricsprecision and recallAnswer metricsrelevancy, correctnessgpt-oss-20bjudge on GroqEmbeddingsGemini
Hover or tap a piece to see what it is and which lesson built it.
Trace a run

Pick one to watch it run, step by step.

A judge with room for long verdicts

The suite builds its own judge from the Groq client in judge.py, with one change: max_tokens=4096. A judge from llm_factory may write about a thousand tokens per reply by default, and gpt-oss-20b spends part of that thinking before it writes. On the longest verdicts of a full run, such as answer correctness on a multi-part answer, it ran out before finishing the JSON and RAGAS raised IncompleteOutputException.

python
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq, max_tokens=4096)  # room for long verdicts

The five experiments

The video's metrics.py keeps a registry: each experiment is a name, a metric and the sample fields it needs. This is that registry with this course's judge and embeddings.

python
EXPERIMENTS = [
    ("faithfulness", Faithfulness(llm=judge), ["user_input", "response", "retrieved_contexts"]),
    ("answer_relevancy", AnswerRelevancy(llm=judge, embeddings=embeddings), ["user_input", "response"]),
    ("context_precision", ContextPrecision(llm=judge), ["user_input", "reference", "retrieved_contexts"]),
    ("context_recall", ContextRecall(llm=judge), ["user_input", "retrieved_contexts", "reference"]),
    ("answer_correctness", AnswerCorrectness(llm=judge, embeddings=embeddings), ["user_input", "response", "reference"]),
]

Phase two: metric by metric

python
async def phase_two(rows, experiments):
    scores = {}
    for n, (name, metric, keys) in enumerate(experiments):
        scores[name] = []
        for i, row in enumerate(rows):
            scores[name].append(await score_one(metric, {k: row[k] for k in keys}))
            if i < len(rows) - 1:
                await asyncio.sleep(SAMPLE_COOLDOWN)
        if n < len(experiments) - 1:
            await asyncio.sleep(EXPERIMENT_COOLDOWN)
    return scores

Each metric gets only the fields its signature names, the list in its registry entry. That is how one loop serves five metrics with different inputs. score_one is the retry from the rate-limits lesson, widened to any error: over a long run a free-tier judge now and then returns a rate limit or an empty reply, and one retry keeps the run from losing ten minutes of scores.

The failure to catch

A retriever change that keeps one chunk instead of three is a realistic regression: someone lowers top_k to save tokens. The script runs phase one again at top_k=1 and scores context recall only, the metric that should catch it.

Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
judge.py
import os

from google import genai
from openai import AsyncOpenAI
from ragas.embeddings import GoogleEmbeddings
from ragas.llms import llm_factory

groq = AsyncOpenAI(
    api_key=os.environ.get("JUDGE_GROQ", os.environ["GROQ_API_KEY"]),
    base_url="https://api.groq.com/openai/v1",
)
judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq)


class OneTextPerCall(GoogleEmbeddings):
    """gemini-embedding-2 turns a list into one embedding, so embed each text on its own."""

    def embed_texts(self, texts, **kwargs):
        return [self.embed_text(text) for text in texts]

    async def aembed_texts(self, texts, **kwargs):
        return [await self.aembed_text(text) for text in texts]


embeddings = OneTextPerCall(client=genai.Client(), model="gemini-embedding-2")
technest.py
import json
import os

import numpy as np
from google import genai
from openai import OpenAI

gemini = genai.Client()  # reads GOOGLE_API_KEY
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
EMBED_MODEL = "gemini-embedding-2"
CHAT_MODEL = "qwen/qwen3.8-27b"

SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""

with open("catalog.json", encoding="utf-8") as f:
    CATALOG = json.load(f)


def embed(texts):
    # gemini-embedding-2 turns everything in one call into one embedding, so send one text per call
    vectors = np.array([gemini.models.embed_content(model=EMBED_MODEL, contents=t).embeddings[0].values for t in texts])
    return vectors / np.linalg.norm(vectors, axis=1, keepdims=True)


DOC_VECTORS = embed([f"{item['title']}. {item['content']}" for item in CATALOG])


def retrieve(question, top_k=3):
    scores = DOC_VECTORS @ embed([question])[0]
    best = np.argsort(scores)[::-1][:top_k]
    return [CATALOG[i]["content"] for i in best]


def generate(question, contexts):
    context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
    ]
    response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content.strip()


def answer(question, top_k=3):
    contexts = retrieve(question, top_k)
    return generate(question, contexts), contexts
catalog.json
[
  {
    "id": "prod_001",
    "category": "product",
    "title": "ProBook X1 Laptop",
    "content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
  },
  {
    "id": "prod_002",
    "category": "product",
    "title": "PixelPhone 15",
    "content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
  },
  {
    "id": "prod_003",
    "category": "product",
    "title": "SoundPods Pro",
    "content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
  },
  {
    "id": "prod_004",
    "category": "product",
    "title": "UltraTab S2 Tablet",
    "content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
  },
  {
    "id": "prod_005",
    "category": "product",
    "title": "SmartWatch X",
    "content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
  },
  {
    "id": "prod_006",
    "category": "product",
    "title": "ProCam 4K Action Camera",
    "content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
  },
  {
    "id": "prod_007",
    "category": "product",
    "title": "BassBuds Max Headphones",
    "content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
  },
  {
    "id": "prod_008",
    "category": "product",
    "title": "SoundBar 360",
    "content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
  },
  {
    "id": "policy_001",
    "category": "policy",
    "title": "Return Policy",
    "content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
  },
  {
    "id": "policy_002",
    "category": "policy",
    "title": "Shipping Policy",
    "content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
  },
  {
    "id": "policy_003",
    "category": "policy",
    "title": "Warranty Policy",
    "content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
  },
  {
    "id": "policy_004",
    "category": "policy",
    "title": "Payment Policy",
    "content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
  },
  {
    "id": "faq_001",
    "category": "faq",
    "title": "Order Tracking",
    "content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
  },
  {
    "id": "faq_002",
    "category": "faq",
    "title": "International Shipping",
    "content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
  },
  {
    "id": "faq_003",
    "category": "faq",
    "title": "Bulk and Business Orders",
    "content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
  }
]
goldens.json
[
  {
    "id": "g001",
    "metric_focus": "faithfulness",
    "user_input": "What is TechNest's return policy?",
    "reference": "TechNest accepts returns within 30 days of purchase. Items must be in original packaging with all accessories. Customers pay return shipping unless the item is defective. Refunds are processed in 5 to 7 business days."
  },
  {
    "id": "g002",
    "metric_focus": "answer_relevancy",
    "user_input": "What are the RAM and storage specs of the ProBook X1?",
    "reference": "The ProBook X1 has 16GB DDR5 RAM and a 512GB NVMe SSD."
  },
  {
    "id": "g003",
    "metric_focus": "context_precision",
    "user_input": "How long is the battery life on the SoundPods Pro?",
    "reference": "The SoundPods Pro offer 8 hours of playback per charge and an additional 24 hours from the charging case, giving a total of 32 hours."
  },
  {
    "id": "g004",
    "metric_focus": "context_recall",
    "user_input": "What are TechNest's shipping options and how long do returns take to process?",
    "reference": "TechNest offers free standard shipping on orders over $50 (3 to 5 business days) and expedited shipping for $9.99 (1 to 2 business days). Returns are accepted within 30 days and refunds are processed in 5 to 7 business days after the item is received."
  },
  {
    "id": "g005",
    "metric_focus": "answer_correctness",
    "user_input": "What is the price of the PixelPhone 15?",
    "reference": "The TechNest PixelPhone 15 is priced at $899."
  }
]

The whole suite in one run

Save this as suite.py in the folder with datasets/goldens.csv, technest.py, catalog.json and judge.py. It scores three of the five goldens, the ones written for faithfulness, context precision and context recall, so the run stays inside Groq's free daily tokens; it still takes several minutes, nearly all of it cooldowns. Put all five ids in SUITE for the video's full run, about fifteen minutes.

ExampleAPI key
import asyncio

from judge import embeddings, groq
from ragas import Dataset, experiment
from ragas.llms import llm_factory
from ragas.metrics.collections import (AnswerCorrectness, AnswerRelevancy, ContextPrecision,
                                       ContextRecall, Faithfulness)
from technest import answer

judge = llm_factory("openai/gpt-oss-20b", provider="openai", client=groq, max_tokens=4096)  # room for long verdicts

SAMPLE_COOLDOWN = 25      # seconds between samples
EXPERIMENT_COOLDOWN = 35  # seconds between metrics
RETRY_WAIT = 65           # seconds to wait before one retry
EXPERIMENTS = [
    ("faithfulness", Faithfulness(llm=judge), ["user_input", "response", "retrieved_contexts"]),
    ("answer_relevancy", AnswerRelevancy(llm=judge, embeddings=embeddings), ["user_input", "response"]),
    ("context_precision", ContextPrecision(llm=judge), ["user_input", "reference", "retrieved_contexts"]),
    ("context_recall", ContextRecall(llm=judge), ["user_input", "retrieved_contexts", "reference"]),
    ("answer_correctness", AnswerCorrectness(llm=judge, embeddings=embeddings), ["user_input", "response", "reference"]),
]


def phase_one(top_k):
    @experiment()
    async def run_app(row):
        response, contexts = answer(row["user_input"], top_k=top_k)
        return {**row, "response": response, "retrieved_contexts": contexts}

    dataset = Dataset.load(name="goldens", backend="local/csv", root_dir=".")
    return sorted(asyncio.run(run_app.arun(dataset, name=f"phase-one-k{top_k}")), key=lambda r: r["id"])


async def score_one(metric, inputs):
    try:
        return (await metric.ascore(**inputs)).value
    except Exception as error:  # a rate limit or a malformed judge reply: wait, then try once more
        print(f"   {type(error).__name__}, retrying once after {RETRY_WAIT}s")
        await asyncio.sleep(RETRY_WAIT)
        return (await metric.ascore(**inputs)).value


async def phase_two(rows, experiments):
    scores = {}
    for n, (name, metric, keys) in enumerate(experiments):
        scores[name] = []
        for i, row in enumerate(rows):
            scores[name].append(await score_one(metric, {k: row[k] for k in keys}))
            if i < len(rows) - 1:
                await asyncio.sleep(SAMPLE_COOLDOWN)
        if n < len(experiments) - 1:
            await asyncio.sleep(EXPERIMENT_COOLDOWN)
    return scores


def badge(score):
    return "good" if score >= 0.75 else "fair" if score >= 0.5 else "poor"


SUITE = ["g001", "g003", "g004"]  # faithfulness, precision and recall goldens; add g002 and g005 for the full run
rows = [r for r in phase_one(top_k=3) if r["id"] in SUITE]
scores = asyncio.run(phase_two(rows, EXPERIMENTS))
names = [name for name, _, _ in EXPERIMENTS]
print("golden  " + "  ".join(f"{n[:12]:<12}" for n in names))
for i, row in enumerate(rows):
    print(f"{row['id']:<8}" + "  ".join(f"{scores[n][i]:.2f} {badge(scores[n][i]):<7}" for n in names))
print("AVERAGE " + "  ".join(f"{sum(scores[n]) / len(rows):.2f} {badge(sum(scores[n]) / len(rows)):<7}" for n in names))

print()
print("Failure mode: the retriever keeps one chunk (top_k=1)")
broken = [r for r in phase_one(top_k=1) if r["id"] in SUITE]
recall_k1 = asyncio.run(phase_two(broken, [EXPERIMENTS[3]]))["context_recall"]
for i, row in enumerate(rows):
    print(f"{row['id']}  context_recall  top_k=3: {scores['context_recall'][i]:.2f}   top_k=1: {recall_k1[i]:.2f}")

What the report shows

  • Faithfulness and context recall are 1.00 on all three goldens. The answers stay inside their chunks, and the chunks hold every claim the references need.
  • g003's answer correctness is 0.53, fair. Faithfulness says every claim in the answer is in the chunk, yet the judge matched the answer's claims to the golden's only in part. This is the cell to read first: set the answer next to the golden and compare claim by claim.
  • g004's context precision is 0.00, poor, while its context recall is 1.00. Recall says every claim of the two-part reference was retrieved; precision says the judge found no single chunk useful for that whole reference. When two metrics disagree this sharply, read the judge's per-chunk verdicts before changing the retriever.
  • The failure run keeps everything the same except top_k. g004, whose reference needs two policies, drops from 1.00 to 0.50; g001 and g003, whose answers sit in a single chunk, keep 1.00. The report points at retrieval, not the prompt.

Other RAGAS metrics and tools

Left outWhat it is
Test set generationBuilding goldens from your documents with an LLM instead of writing them
Multi-turn metrics and MultiTurnSampleGrading a whole conversation rather than one answer
Agent metricsTool call accuracy, agent goal accuracy and topic adherence
Noise sensitivity, context entity recall, response groundednessMore single-turn RAG metrics
Traditional metricsBLEU, ROUGE, CHRF and string distances, some needing extra packages
Rubrics and aspect criticsScoring against written rubrics, a step beyond DiscreteMetric
Metric prompt alignmentTuning a metric's prompt against your own labelled examples
Framework integrationsEvaluating LangChain, LlamaIndex and other apps from their own traces

Extending the suite

  • Add goldens from real support tickets, starting with every answer a user complained about.
  • Split the goldens into essential and full sets and put the essential ones in CI, as in Evals in CI.
  • Swap the judge to another provider, as in Integrations, and re-run the baseline before comparing.
Watch out. Do not compare this report with the video's numbers. The video judged with a different model on a different day; only runs of the same suite, same judge and same goldens compare with each other.
Try it yourself
  • Break faithfulness instead: add "Return shipping is free on every order." to every response in phase_one and see which metrics fall.
  • Add the tone DiscreteMetric from the DiscreteMetric lesson as a sixth experiment.
  • Lower SAMPLE_COOLDOWN to 10 and watch for rate-limit errors on your key.
PreviousIntegrations

You understood something today that you didn't yesterday.