Enterprise RAG assistant: cited answers and an eval dashboard
The month 2 project of the AI Forward Deployed Engineer roadmap: an assistant that answers from a corpus of policies or contracts, cites the document and page behind every answer, refuses when the corpus is silent, and comes with an eval dashboard that shows how much your tuning improved it.
The problem
Staff answer questions from long policy and contract documents, and a wrong answer given confidently is worse than no answer. What makes an answer usable is a citation: which document, which section, which page, so a person can check it before acting on it.
A demo can do that on five questions. The client needs to know it does it on the questions their staff actually ask, and that the next change will not quietly make it worse. So the deliverable is two things: the assistant, and the numbers that prove it works.
Architecture
A retrieval pipeline that keeps the source at every step, with an evaluation loop beside it. Documents are parsed and chunked so each passage remembers its document and page, indexed for both meaning and exact words, retrieved and reranked for a question, and answered only from what came back. Below it, a golden set of real questions runs through the same pipeline and RAGAS scores every answer, so each tuning step shows up on the dashboard.
Measure the baseline before you tune anything. Then change one thing at a time: chunk size, hybrid search, a reranker, the similarity cutoff, the prompt. Score after each change, so the dashboard tells a story the client can follow.
What it draws on
- LangChain: documents, embeddings and a vector store
- LlamaIndex: chunking, metadata filters, hybrid search, reranking and hit rate
- LlamaIndex: answers with sources and refusing
- LangGraph: grounded RAG in a graph
- Topic: golden datasets
- RAGAS: LLM-as-judge and the core RAG metrics
- Vector databases (pgvector, Qdrant, Pinecone), query rewriting, HyDE and PageIndex: courses upcoming. Use their official docs if you want them in the project
What done looks like
| Requirement | Done when |
|---|---|
| Corpus | At least 20 real or realistic policy or contract documents, indexed with document, section and page |
| Citations | Every answer names the document and page of each claim, and the passage really says it |
| Refusal | A question the corpus does not cover gets a clear refusal, not a guess |
| Golden set | At least 30 questions with agreed answers, including ones the corpus cannot answer |
| Scores | RAGAS faithfulness, answer relevancy, context precision and recall, and correctness for every golden |
| Dashboard | A page or notebook that shows each metric for the baseline and after each tuning step |
Where to start
Get one document answering one question with a real citation, and write your first ten goldens on the same day. Score the baseline. Only then add documents, hybrid search and a reranker, scoring after each one. If a change does not move a number, take it out again.
Every expert started right here.