Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q10EasyConcept

How do you detect and handle PII in LLM applications?

30-second answerSay your answer out loud first, then reveal.

Techniques

MethodGood forLimitations
Regex + validation (Luhn for cards, format checks)Structured IDsMisses unstructured PII; false positives
NER models (spaCy, transformer NER, Presidio recognisers)Names, locations, organisationsLanguage and domain sensitivity
LLM-based detectionContext-dependent PII ("my neighbour's diagnosis")Slower, costlier, non-deterministic
Allow / deny listsCompany names, product termsMaintenance

Handling strategies

  • Masking: simplest, but loses meaning.
  • Pseudonymisation: keeps relationships for the model; restore after generation if needed.
  • Synthetic replacement: realistic fake values (for test data).
  • Blocking: for prohibited data (e.g. full card numbers in chat).

Evaluation: a labelled PII test set (including Indian names in Latin and Devanagari scripts, addresses, mixed-language messages); track recall (missed PII is a privacy risk) and precision (over-masking hurts usefulness).

Don't forget logs and traces: often the largest PII exposure.

Every expert started right here.