Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q10EasyConcept

How do metadata and filtering improve retrieval?

30-second answerSay your answer out loud first, then reveal.

Examples

  • "What's the leave policy?" from an employee in India → filter country = IN, doc_type = policy, status = current.
  • "What changed in v2.3 of the SDK?" → filter version = 2.3.
  • Multi-tenant SaaS → always filter tenant_id = X (security, not just relevance).

Where metadata comes from

  • Source systems (file path, author, modified date, space or folder).
  • Parsing (section headings, page numbers).
  • LLM extraction at index time (topics, entities, document type).

Self-querying retrieval: an LLM converts "Show me Q3 2026 sales reports from the APAC team" into {query: "sales report", filters: {quarter: "Q3-2026", team: "APAC"}}. Validate the extracted filters against allowed values, since the LLM can invent fields.

Pre-filter vs post-filter

  • Pre-filter (filter, then search): correct results, but can be slow or reduce ANN recall with very selective filters.
  • Post-filter (search, then filter): fast, but can return fewer than k results if most hits get filtered out.
  • Modern vector DBs support filtered ANN that handles this in the engine.

Common mistakes

  • Using metadata filters as the only security mechanism without testing them (see Q37).
  • Over-filtering from wrongly extracted filters, which silently returns nothing.

Every expert started right here.