Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q26IntermediateConcept

What does operating the RAG data pipeline involve in production?

30-second answerSay your answer out loud first, then reveal.

Operational checklist

AreaPractice
IngestionConnectors with retries, dead-letter queues, idempotent upserts by doc ID
FreshnessSLA per source (e.g. Confluence ≤ 1 hour); lag dashboards; alerts on stalled syncs
Deletes and ACLsDeletes and permission revocations prioritised; reconciliation jobs comparing source vs index
Quality checksParse-quality heuristics (empty text, garbled characters); sample reviews of new sources
VersioningIndex versions with metadata (embedder, chunker, date); alias switch; keep previous for rollback
EvalsRetrieval recall/MRR on golden queries after each rebuild or config change
BackfillsRe-embedding jobs with checkpointing, throttling to protect APIs, cost estimates
MonitoringDoc counts per source, chunk counts, top-1 similarity distribution, no-result rate

Common incident: a connector silently fails (expired token), and answers go stale for weeks. Freshness alerts would have caught it on day one.

Little by little, you're building something great.