1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
The answer exists in your documents, but the retriever never returns it. How do you debug?
30-second answerSay your answer out loud first, then reveal.
Step-by-step checklist
- Is it in the index? Search the raw chunk store for a unique phrase from the expected answer.
• Not found → parsing or ingestion bug (failed file, OCR gap, filtered out as boilerplate, not synced yet). - Is the chunk self-contained? The answer may be in a chunk that says "It is 30 days" with "the refund window" in the previous chunk. Fix with contextual headers (Q22), larger chunks, or parent-child retrieval (Q21).
- Where does it rank? Compute the similarity between the query and the chunk, and its rank among all chunks.
• Rank 12 with k=5 → reranking, bigger candidate pool.
• Rank 3,000 → semantic mismatch. - Semantic mismatch causes:
• Acronyms or jargon ("LOP" vs "loss of pay") → BM25 hybrid, synonym expansion, query rewriting.
• Exact IDs → BM25.
• Domain language → fine-tuned embeddings.
• Question vs statement style → HyDE, or index generated questions per chunk ("hypothetical questions"). - Filters: is a metadata or permission filter excluding it (wrong date, wrong department tag)?
- Embedding usage bugs: missing
query:/passage:prefixes, different models used for indexing and querying, text truncated beyond the token limit, normalisation mismatch. - Add the case to the retrieval eval set and track recall@k after the fix.
Interview signal. Methodical, stage-by-stage isolation, not "I'd try a bigger LLM."
Related
Slow is fine. Stopping is the only problem.