Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q29IntermediateConcept

What techniques defend against prompt injection? Which ones actually work?

30-second answerSay your answer out loud first, then reveal.
Untrusted content passes through spotlighting and an injection classifier into a quarantined LLM that returns only structured fields to a privileged agent, which also takes the trusted user request and reaches tools and egress only through a policy layer.

Effectiveness ranking (practical view)

DefenceStrengthNotes
Least privilege / no dangerous tools in contexts with untrusted dataStrongRemoves the impact even if injection succeeds
Human approval for sensitive actionsStrongWatch for approval fatigue
Dual-LLM / quarantine, capability or data-flow tracking (CaMeL-style)Strong but complexUntrusted text never directly drives tool arguments
Egress controls (block markdown image exfiltration, URL allow-lists)Strong for exfiltrationCheap and important
Instruction hierarchy in modelsMediumModels trained to prioritise system > user > tool content
Spotlighting / delimiters / "treat as data"Medium-weakReduces success rates; bypassable
Injection classifiersMedium-weakUseful signal; adversaries adapt
Key interview line. "I assume injection will sometimes succeed, and design so that a successful injection can't do serious damage."

This is what real progress feels like.