RAGASragas 0.4.3 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

TechNest RAG app

A RAG app is a program that retrieves the knowledge-base chunks closest to a question and passes them to an LLM, which writes the answer from those chunks only.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

Evaluation needs something to evaluate. This lesson builds the TechNest support bot from the video's repository as one file, technest.py, so every later lesson can ask it questions and score what it says.

The TechNest app and its knowledge base · from the Production RAG Live Marathon · 164:03 to 167:05

The video's evaluation app takes three keys: a Groq key for the answers, a judge key, and a Gemini key for the embeddings. Its knowledge base belongs to a store that sells tech devices online. Each entry has an id, a category, a title and the content: products, shipping and return policies, and order tracking.

The retrieve and generate API

python
contexts = retrieve(question, top_k=3)   # the 3 catalog entries closest to the question
response = generate(question, contexts)  # the LLM answers from those entries only
response, contexts = answer(question)    # both steps together

The knowledge base

Save the video's catalog as catalog.json: 8 products, 4 policies and 3 FAQ entries. This is the file the app searches.

json
[
  {
    "id": "prod_001",
    "category": "product",
    "title": "ProBook X1 Laptop",
    "content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
  },
  {
    "id": "prod_002",
    "category": "product",
    "title": "PixelPhone 15",
    "content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
  },
  {
    "id": "prod_003",
    "category": "product",
    "title": "SoundPods Pro",
    "content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
  },
  {
    "id": "prod_004",
    "category": "product",
    "title": "UltraTab S2 Tablet",
    "content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
  },
  {
    "id": "prod_005",
    "category": "product",
    "title": "SmartWatch X",
    "content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
  },
  {
    "id": "prod_006",
    "category": "product",
    "title": "ProCam 4K Action Camera",
    "content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
  },
  {
    "id": "prod_007",
    "category": "product",
    "title": "BassBuds Max Headphones",
    "content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
  },
  {
    "id": "prod_008",
    "category": "product",
    "title": "SoundBar 360",
    "content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
  },
  {
    "id": "policy_001",
    "category": "policy",
    "title": "Return Policy",
    "content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
  },
  {
    "id": "policy_002",
    "category": "policy",
    "title": "Shipping Policy",
    "content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
  },
  {
    "id": "policy_003",
    "category": "policy",
    "title": "Warranty Policy",
    "content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
  },
  {
    "id": "policy_004",
    "category": "policy",
    "title": "Payment Policy",
    "content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
  },
  {
    "id": "faq_001",
    "category": "faq",
    "title": "Order Tracking",
    "content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
  },
  {
    "id": "faq_002",
    "category": "faq",
    "title": "International Shipping",
    "content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
  },
  {
    "id": "faq_003",
    "category": "faq",
    "title": "Bulk and Business Orders",
    "content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
  }
]

Embedding the catalog with Gemini

An embedding turns a text into a list of numbers, so that texts with similar meaning get similar lists. embed sends each text to Gemini's gemini-embedding-2 in its own call, because that model merges everything sent in one call into a single embedding. Then it scales each list to length 1, which makes comparing two of them a single multiplication.

python
gemini = genai.Client()  # reads GOOGLE_API_KEY
EMBED_MODEL = "gemini-embedding-2"


def embed(texts):
    # gemini-embedding-2 turns everything in one call into one embedding, so send one text per call
    vectors = np.array([gemini.models.embed_content(model=EMBED_MODEL, contents=t).embeddings[0].values for t in texts])
    return vectors / np.linalg.norm(vectors, axis=1, keepdims=True)


DOC_VECTORS = embed([f"{item['title']}. {item['content']}" for item in CATALOG])

The catalog is embedded once, when the file is imported. Each entry is embedded as its title plus its content, as the video's retriever does.

Retrieving the top three chunks

python
def retrieve(question, top_k=3):
    scores = DOC_VECTORS @ embed([question])[0]
    best = np.argsort(scores)[::-1][:top_k]
    return [CATALOG[i]["content"] for i in best]

@ multiplies the question's vector with every catalog vector at once, giving one similarity score per entry. The three highest win, best first. That order matters later: context precision scores it.

Generating the answer with Groq

The system prompt is the video's, word for word, plus one last line that keeps answers short.

python
def generate(question, contexts):
    context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
    ]
    response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content.strip()

The numbered chunks go in the user message, and the prompt tells the model to answer from them only and to say so when they are not enough.

The whole technest.py file

Put the pieces together in technest.py, next to catalog.json. Later lessons import answer and retrieve from it.

python
import json
import os

import numpy as np
from google import genai
from openai import OpenAI

gemini = genai.Client()  # reads GOOGLE_API_KEY
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
EMBED_MODEL = "gemini-embedding-2"
CHAT_MODEL = "qwen/qwen3.8-27b"

SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""

with open("catalog.json", encoding="utf-8") as f:
    CATALOG = json.load(f)


def embed(texts):
    # gemini-embedding-2 turns everything in one call into one embedding, so send one text per call
    vectors = np.array([gemini.models.embed_content(model=EMBED_MODEL, contents=t).embeddings[0].values for t in texts])
    return vectors / np.linalg.norm(vectors, axis=1, keepdims=True)


DOC_VECTORS = embed([f"{item['title']}. {item['content']}" for item in CATALOG])


def retrieve(question, top_k=3):
    scores = DOC_VECTORS @ embed([question])[0]
    best = np.argsort(scores)[::-1][:top_k]
    return [CATALOG[i]["content"] for i in best]


def generate(question, contexts):
    context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
    ]
    response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content.strip()


def answer(question, top_k=3):
    contexts = retrieve(question, top_k)
    return generate(question, contexts), contexts

This file differs from the video's repository in four places. The repository embeds with gemini-embedding-2-preview; this file uses gemini-embedding-2, the released version of that model. The repository answers with llama-3.3-70b-versatile, which Groq has retired from its free plan; this file uses qwen/qwen3.8-27b, a different model family from the judge the course builds later. The repository wraps the calls in async classes for its Streamlit app; a script needs only plain functions. And the extra prompt line keeps each answer short, so the judge reads fewer tokens per call.

Asking the bot the first golden question

ExampleAPI keyFrom the video's repository, run on Groq and Gemini
from technest import answer

response, contexts = answer("What is TechNest's return policy?")
for number, chunk in enumerate(contexts, 1):
    print(f"chunk {number}: {chunk[:70]}...")
print()
print(response)

What the bot retrieved and said

  • chunk 1 is the return policy, the entry the question needs, ranked first.
  • chunks 2 and 3 are the warranty policy and the shipping policy, the next closest entries. They do not answer the question; whether they hurt is what the retrieval metrics measure.
  • The answer is written from the chunks. Whether every sentence in it is backed by a chunk is what faithfulness checks.

A plain RAG app vs one with evaluation

Plain RAG appRAG app with evaluation
What you seeOne answer at a timeA score per question and per quality
Which chunks were usedHiddenStored with the answer
When the answer is wrongYou find out from usersA low score shows it before release

When to build the app this way

  • When you want retrieval and generation as separate functions, so an evaluation can store what each returned.
  • When the knowledge base is small enough to hold in memory. For a large one, the Integrations lesson swaps the list for a vector store.
Watch out. Keep the retrieved chunks, not only the answer. An app that returns only the text makes it impossible to tell later whether a wrong answer came from the search or from the model.
Try it yourself
  • Ask answer("How long is the battery life on the SoundPods Pro?") and read which chunks come back.
  • Call retrieve("Do you ship to Canada?", top_k=5) and check whether the international shipping FAQ is first.
  • Ask something the catalog does not cover, such as the price of a TV, and read how the prompt makes the bot answer.

Slow is fine. Stopping is the only problem.