RAGAS overview
RAGAS is an open-source Python library that scores a retrieval-augmented generation (RAG) app, using an LLM as a judge to check whether each answer stays true to the retrieved documents and whether the search found the right documents.
Last updated: 29 Sep, 2026 · RAGAS 0.4.3
RAGAS stands for Retrieval Augmented Generation Assessment. It started as a 2023 research paper and grew into a general library for evaluating LLM applications. It is released under the Apache 2.0 licence, developed on GitHub by Vibrant Labs, and this course uses version 0.4.3.
A RAG app answers in two steps. It searches a knowledge base for the chunks that match the question, then an LLM writes the answer from those chunks. When an answer is wrong, either the search missed or the writing drifted. RAGAS gives each half its own scores, so a low number points at the part to fix.
The video opens on a plain RAG chatbot. It uploads the Attention Is All You Need paper, asks about self-attention, and gets an answer. The answer looks fine, but nothing on the screen says how good it is, which chunks it came from, or how the bot would handle the many other questions real users will ask. That missing view is what an evaluation pipeline adds.
Evaluating the TechNest support bot
The video then opens a second app: the same kind of RAG bot, this time for TechNest, an online electronics store, with an evaluation pipeline around it. This course builds that app and that pipeline in Python, one piece per lesson.
- The app. A catalog of 15 products and policies, a retriever that finds the three closest entries with Gemini embeddings, and a Groq model that writes the answer.
- The goldens. Five questions a customer might ask, each with the answer TechNest expects.
- The judge. A second Groq model that reads each answer and scores it with five RAGAS metrics: faithfulness, answer relevancy, context precision, context recall and answer correctness.
- The report. A score per question and per metric, so you can see which question fails and on which side.
Pick one to watch it run, step by step.
Samples, metrics, a judge and experiments
These are the RAGAS names you will meet, in the order the course introduces them.
| RAGAS piece | What it holds or does | Taught in |
|---|---|---|
SingleTurnSample | One question, the chunks that came back, the answer, and the expected answer | SingleTurnSample |
| Metrics with no model | Plain Python scores: exact match, string presence, id overlap | Exact match and string presence |
llm_factory | Wraps a hosted model so it can act as the judge | LLM as a judge |
| The five RAG metrics | Faithfulness, answer relevancy, context precision, context recall, answer correctness | RAGAS metrics |
Dataset and @experiment | Goldens in a file, and one saved run of the app over them | Experiments |
evaluate() | The older one-call API most existing examples use | evaluate() and EvaluationDataset |
Keys you need
Every score in this course comes from a real model, so you need two free keys. Neither needs a card.
| Key | Used for | Where to get it |
|---|---|---|
GROQ_API_KEY | The app's answers and the judge | console.groq.com/keys |
GOOGLE_API_KEY | Gemini embeddings: the app's search, and two metrics | aistudio.google.com/apikey |
Related
- Next: Installation and setup
- Reference: RAGAS documentation
Every expert started right here.