Dashboard

Enterprise RAG assistant: cited answers and an eval dashboard

The month 2 project of the AI Forward Deployed Engineer roadmap: an assistant that answers from a corpus of policies or contracts, cites the document and page behind every answer, refuses when the corpus is silent, and comes with an eval dashboard that shows how much your tuning improved it.

The problem

Staff answer questions from long policy and contract documents, and a wrong answer given confidently is worse than no answer. What makes an answer usable is a citation: which document, which section, which page, so a person can check it before acting on it.

A demo can do that on five questions. The client needs to know it does it on the questions their staff actually ask, and that the next change will not quietly make it worse. So the deliverable is two things: the assistant, and the numbers that prove it works.

Architecture

A retrieval pipeline that keeps the source at every step, with an evaluation loop beside it. Documents are parsed and chunked so each passage remembers its document and page, indexed for both meaning and exact words, retrieved and reranked for a question, and answered only from what came back. Below it, a golden set of real questions runs through the same pipeline and RAGAS scores every answer, so each tuning step shows up on the dashboard.

A policy corpus in, cited answers out, and a score that shows tuning helped
Policy corpuscontracts, PDFsParse and chunksource + page keptHybrid indexvectors + BM25Retrieve, reranktop passagesCited answeror a refusalGolden setreal questionsRAGAS scoresscored by a judgeEval dashboardbefore vs after
Hover or tap a piece to see what it does.

Measure the baseline before you tune anything. Then change one thing at a time: chunk size, hybrid search, a reranker, the similarity cutoff, the prompt. Score after each change, so the dashboard tells a story the client can follow.

What it draws on

What done looks like

RequirementDone when
CorpusAt least 20 real or realistic policy or contract documents, indexed with document, section and page
CitationsEvery answer names the document and page of each claim, and the passage really says it
RefusalA question the corpus does not cover gets a clear refusal, not a guess
Golden setAt least 30 questions with agreed answers, including ones the corpus cannot answer
ScoresRAGAS faithfulness, answer relevancy, context precision and recall, and correctness for every golden
DashboardA page or notebook that shows each metric for the baseline and after each tuning step

Where to start

Get one document answering one question with a real citation, and write your first ten goldens on the same day. Score the baseline. Only then add documents, hybrid search and a reranker, scoring after each one. If a change does not move a number, take it out again.

A free Groq key covers the answering model and the RAGAS judge, and a free Gemini key covers the embeddings RAGAS uses, as the course setup lessons show. The LlamaIndex lessons embed with a small local model and need no key. Keep keys in the environment, never in the repository.
Back toAll frameworks

Every expert started right here.