1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your path
A stakeholder says "the RAG bot's accuracy is bad." You have two weeks. How do you approach it end to end?
30-second answerSay your answer out loud first, then reveal.
Days 1–2: define and collect
- Meet stakeholders: who uses it, for what, which failures hurt most? Get 30–50 real bad examples.
- Agree on success metrics (e.g. correct answer rate on a golden set, abstention when the answer isn't present, citation accuracy).
Days 3–4: build an eval set and baseline
- 100–200 questions with reference answers and source documents (real + synthetic + unanswerable).
- Measure the baseline: retrieval recall@k, faithfulness, correctness.
Days 5–6: error analysis
- For each failure, find the first broken stage (Q14). Example distribution:
| Category | Share | Likely fix |
|---|---|---|
| Content missing from KB | 20% | Content gap report to owners; abstain properly |
| Parsing (tables, PDFs) | 15% | Better parser; table handling |
| Retrieval miss (acronyms, IDs) | 25% | Hybrid BM25 + query rewriting |
| Ranking (relevant at rank 8–30) | 15% | Reranker; tune k |
| Generation (ignored context, hallucinated) | 15% | Prompt grounding, citations, model |
| Outdated / conflicting docs | 10% | Version metadata; dedupe |
Days 7–11: fix in order of impact ÷ effort
Re-run the eval after each change, keep what helps, revert what doesn't.
Days 12–14: ship and institutionalise
- Canary release; dashboard (feedback, abstention, latency); a weekly error-review ritual; the eval set in CI.
- Report: baseline → current per metric, remaining gaps, next steps.
What interviewers are testing. Structured thinking, measurement before changes, prioritisation, and stakeholder communication. That's the core skill for AI engineers and FDEs.
Related
Every expert started right here.