1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you handle privacy and compliance in RAG: PII, the right to be forgotten, and multi-tenancy?
30-second answerSay your answer out loud first, then reveal.
Key areas
- PII at ingestion:
◦ Detect it with NER-based tools (e.g. Microsoft Presidio) plus regex for IDs (Aadhaar, PAN, phone, email).
◦ Policy per source: exclude, redact, or keep with stricter ACLs. - Embeddings are not anonymous: research has shown that text can be partially reconstructed from embeddings. Protect vector stores like the source data.
- Right to be forgotten / deletion:
◦ Maintain a lineage map: source doc → chunks → vectors → cached answers → derived summaries.
◦ Delete across all of them within the SLA, then verify.
◦ Don't forget logs, traces and eval datasets containing user data. - Multi-tenancy:
◦ Tenant ID enforced at the storage and query layer (namespaces, separate collections, row-level security).
◦ Separate encryption keys per tenant for strict customers.
◦ Caches keyed by tenant and permission scope.
◦ Automated cross-tenant leakage tests. - LLM provider considerations: data residency (e.g. India or EU regions), no-training-on-data agreements, retention settings; self-hosted models for the most sensitive workloads.
- Output controls: a PII filter on answers; prevent the model from revealing other users' data from conversation memory.
- Audit: who asked what, which documents were retrieved, and what was shown.
Interview line. "Every copy of the data the RAG pipeline creates (chunks, vectors, caches, logs) inherits the compliance obligations of the original."
Related
This is what real progress feels like.