1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do metadata and filtering improve retrieval?
30-second answerSay your answer out loud first, then reveal.
Examples
- "What's the leave policy?" from an employee in India → filter
country = IN, doc_type = policy, status = current. - "What changed in v2.3 of the SDK?" → filter
version = 2.3. - Multi-tenant SaaS → always filter
tenant_id = X(security, not just relevance).
Where metadata comes from
- Source systems (file path, author, modified date, space or folder).
- Parsing (section headings, page numbers).
- LLM extraction at index time (topics, entities, document type).
Self-querying retrieval: an LLM converts "Show me Q3 2026 sales reports from the APAC team" into {query: "sales report", filters: {quarter: "Q3-2026", team: "APAC"}}. Validate the extracted filters against allowed values, since the LLM can invent fields.
Pre-filter vs post-filter
- Pre-filter (filter, then search): correct results, but can be slow or reduce ANN recall with very selective filters.
- Post-filter (search, then filter): fast, but can return fewer than k results if most hits get filtered out.
- Modern vector DBs support filtered ANN that handles this in the engine.
Common mistakes
- Using metadata filters as the only security mechanism without testing them (see Q37).
- Over-filtering from wrongly extracted filters, which silently returns nothing.
Related
Every expert started right here.