RAGASragas 0.4.3 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
28 small wins to finish your pathNext lesson →

RAGAS logoRAGAS overview

RAGAS is an open-source Python library that scores a retrieval-augmented generation (RAG) app, using an LLM as a judge to check whether each answer stays true to the retrieved documents and whether the search found the right documents.

Last updated: 29 Sep, 2026 · RAGAS 0.4.3

RAGAS stands for Retrieval Augmented Generation Assessment. It started as a 2023 research paper and grew into a general library for evaluating LLM applications. It is released under the Apache 2.0 licence, developed on GitHub by Vibrant Labs, and this course uses version 0.4.3.

A RAG app answers in two steps. It searches a knowledge base for the chunks that match the question, then an LLM writes the answer from those chunks. When an answer is wrong, either the search missed or the writing drifted. RAGAS gives each half its own scores, so a low number points at the part to fix.

A RAG chatbot with no evaluation · from the Production RAG Live Marathon · 160:48 to 164:03

The video opens on a plain RAG chatbot. It uploads the Attention Is All You Need paper, asks about self-attention, and gets an answer. The answer looks fine, but nothing on the screen says how good it is, which chunks it came from, or how the bot would handle the many other questions real users will ask. That missing view is what an evaluation pipeline adds.

Evaluating the TechNest support bot

The video then opens a second app: the same kind of RAG bot, this time for TechNest, an online electronics store, with an evaluation pipeline around it. This course builds that app and that pipeline in Python, one piece per lesson.

  • The app. A catalog of 15 products and policies, a retriever that finds the three closest entries with Gemini embeddings, and a Groq model that writes the answer.
  • The goldens. Five questions a customer might ask, each with the answer TechNest expects.
  • The judge. A second Groq model that reads each answer and scores it with five RAGAS metrics: faithfulness, answer relevancy, context precision, context recall and answer correctness.
  • The report. A score per question and per metric, so you can see which question fails and on which side.
The evaluation suite you will build, and the lesson each piece comes from
five metrics, one at a timerowsquestionanswer + chunksascore()judgesembedswritesDatasetgoldens.csvlesson 19@experimentrun_app(row)lesson 20technest.answerresponse, chunkslesson 3phase-one-k3.csvexperiments/lesson 20Faithfulnessclaims in the chunks?lesson 11Retrieval metricsprecision and recalllesson 13Answer metricsrelevancy, correctnesslesson 12gpt-oss-20bjudge on Groqlesson 8EmbeddingsGeminilesson 8
Hover or tap a piece to see what it is and which lesson built it.
Trace a run

Pick one to watch it run, step by step.

Samples, metrics, a judge and experiments

These are the RAGAS names you will meet, in the order the course introduces them.

RAGAS pieceWhat it holds or doesTaught in
SingleTurnSampleOne question, the chunks that came back, the answer, and the expected answerSingleTurnSample
Metrics with no modelPlain Python scores: exact match, string presence, id overlapExact match and string presence
llm_factoryWraps a hosted model so it can act as the judgeLLM as a judge
The five RAG metricsFaithfulness, answer relevancy, context precision, context recall, answer correctnessRAGAS metrics
Dataset and @experimentGoldens in a file, and one saved run of the app over themExperiments
evaluate()The older one-call API most existing examples useevaluate() and EvaluationDataset

Keys you need

Every score in this course comes from a real model, so you need two free keys. Neither needs a card.

KeyUsed forWhere to get it
GROQ_API_KEYThe app's answers and the judgeconsole.groq.com/keys
GOOGLE_API_KEYGemini embeddings: the app's search, and two metricsaistudio.google.com/apikey
Watch out. A judged score changes a little from run to run, because the judge is a model. The scores on these pages are real runs; yours will be close to them, not identical. Read them as a direction, not as the last decimal.
Back toAll frameworks

Every expert started right here.