Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q31IntermediateConcept

How do you protect sensitive data (PII) when using third-party LLM APIs?

30-second answerSay your answer out loud first, then reveal.
PII protection flow: input goes through PII detection and a sensitivity check; restricted data goes to a self-hosted model, confidential data is pseudonymised using an encrypted short-TTL token map, sent to a zero-retention regional external LLM API, and re-identified in the response.

Considerations

  1. Detection quality: Indian identifiers (Aadhaar, PAN, GSTIN, IFSC), names in multiple scripts, free-form addresses. Combine regex, NER models and allow/deny lists, and evaluate recall.
  2. Pseudonymisation vs masking: consistent placeholders preserve meaning ("PERSON_1 called PERSON_2"), so the model can still reason. Pure masking ([REDACTED]) loses relationships.
  3. Context leakage: retrieved documents and tool outputs also contain PII, so apply the same pipeline to them.
  4. Logs and traces: often the biggest leak. Redact before logging; restrict access; apply retention limits.
  5. Contracts and settings: data processing agreements, no-training clauses, retention settings, data residency (e.g. India region for regulated sectors).
  6. Output checks: ensure responses don't expose other users' data.

Little by little, you're building something great.