1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you detect and handle PII in LLM applications?
30-second answerSay your answer out loud first, then reveal.
Techniques
| Method | Good for | Limitations |
|---|---|---|
| Regex + validation (Luhn for cards, format checks) | Structured IDs | Misses unstructured PII; false positives |
| NER models (spaCy, transformer NER, Presidio recognisers) | Names, locations, organisations | Language and domain sensitivity |
| LLM-based detection | Context-dependent PII ("my neighbour's diagnosis") | Slower, costlier, non-deterministic |
| Allow / deny lists | Company names, product terms | Maintenance |
Handling strategies
- Masking: simplest, but loses meaning.
- Pseudonymisation: keeps relationships for the model; restore after generation if needed.
- Synthetic replacement: realistic fake values (for test data).
- Blocking: for prohibited data (e.g. full card numbers in chat).
Evaluation: a labelled PII test set (including Indian names in Latin and Devanagari scripts, addresses, mixed-language messages); track recall (missed PII is a privacy risk) and precision (over-masking hurts usefulness).
Don't forget logs and traces: often the largest PII exposure.
Related
Every expert started right here.