Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q46HardConcept

How do you monitor a RAG system in production and improve it continuously?

30-second answerSay your answer out loud first, then reveal.

What to log per request

  • Raw and transformed queries, filters, user segment (not raw PII where avoidable).
  • Retrieved chunk IDs, scores, ranks (pre- and post-rerank).
  • Prompt version, model, token counts, latency per stage, cost.
  • Answer, citations, and the user's feedback.

Dashboards and alerts

CategoryMetrics
Quality (online)Thumbs up/down, follow-up rephrasing rate ("that's not what I asked"), escalation to human, citation click-through
Quality (sampled judge)Faithfulness, answer relevance on X% of traffic
Retrieval healthTop-1 score distribution (drops signal drift), no-results rate, % answers with zero citations
Content gapsClusters of unanswered or low-score questions → missing documentation
Systemp50/p95 latency per stage, error rates, cost per query
FreshnessIndex lag per source, failed syncs

Continuous improvement loop

  1. Weekly: sample failing and negative-feedback traces → categorise (retrieval miss, parsing, generation, content gap).
  2. Fix the top category; add the cases to the eval set.
  3. Run the offline eval → ship → monitor.
  4. Share content gap reports with documentation owners. Often the best fix is writing the missing doc.

Drift: new products, new terminology, seasonal queries. Clustering query embeddings over time helps spot new topics.

Little by little, you're building something great.