Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q18IntermediateScenario

The answer exists in your documents, but the retriever never returns it. How do you debug?

30-second answerSay your answer out loud first, then reveal.

Step-by-step checklist

  1. Is it in the index? Search the raw chunk store for a unique phrase from the expected answer.
    • Not found → parsing or ingestion bug (failed file, OCR gap, filtered out as boilerplate, not synced yet).
  2. Is the chunk self-contained? The answer may be in a chunk that says "It is 30 days" with "the refund window" in the previous chunk. Fix with contextual headers (Q22), larger chunks, or parent-child retrieval (Q21).
  3. Where does it rank? Compute the similarity between the query and the chunk, and its rank among all chunks.
    • Rank 12 with k=5 → reranking, bigger candidate pool.
    • Rank 3,000 → semantic mismatch.
  4. Semantic mismatch causes:
    • Acronyms or jargon ("LOP" vs "loss of pay") → BM25 hybrid, synonym expansion, query rewriting.
    • Exact IDs → BM25.
    • Domain language → fine-tuned embeddings.
    • Question vs statement style → HyDE, or index generated questions per chunk ("hypothetical questions").
  5. Filters: is a metadata or permission filter excluding it (wrong date, wrong department tag)?
  6. Embedding usage bugs: missing query: / passage: prefixes, different models used for indexing and querying, text truncated beyond the token limit, normalisation mismatch.
  7. Add the case to the retrieval eval set and track recall@k after the fix.
Interview signal. Methodical, stage-by-stage isolation, not "I'd try a bigger LLM."

Slow is fine. Stopping is the only problem.