1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you monitor a RAG system in production and improve it continuously?
30-second answerSay your answer out loud first, then reveal.
What to log per request
- Raw and transformed queries, filters, user segment (not raw PII where avoidable).
- Retrieved chunk IDs, scores, ranks (pre- and post-rerank).
- Prompt version, model, token counts, latency per stage, cost.
- Answer, citations, and the user's feedback.
Dashboards and alerts
| Category | Metrics |
|---|---|
| Quality (online) | Thumbs up/down, follow-up rephrasing rate ("that's not what I asked"), escalation to human, citation click-through |
| Quality (sampled judge) | Faithfulness, answer relevance on X% of traffic |
| Retrieval health | Top-1 score distribution (drops signal drift), no-results rate, % answers with zero citations |
| Content gaps | Clusters of unanswered or low-score questions → missing documentation |
| System | p50/p95 latency per stage, error rates, cost per query |
| Freshness | Index lag per source, failed syncs |
Continuous improvement loop
- Weekly: sample failing and negative-feedback traces → categorise (retrieval miss, parsing, generation, content gap).
- Fix the top category; add the cases to the eval set.
- Run the offline eval → ship → monitor.
- Share content gap reports with documentation owners. Often the best fix is writing the missing doc.
Drift: new products, new terminology, seasonal queries. Clustering query embeddings over time helps spot new topics.
Related
- Previous: Q45. You switched to a new embedding model and a new chunking strategy, and production quality dropped. How should you have managed this migration, and what now?
- Next: Q47. Users ask "How many of our vendor contracts have an auto-renewal clause?" and RAG gives wrong answers. Why, and what would you build?
- RAGAS
- DeepEval
Little by little, you're building something great.