1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is prompt injection in the context of agents? Direct vs indirect?
30-second answerSay your answer out loud first, then reveal.

Why it's hard: LLMs can't reliably separate "instructions" from "data". Everything is tokens in the same context. There is currently no complete fix at the model level, so defence relies on system design.
Defences (defence in depth)
- Least privilege: give agents only the tools and scopes they need. An agent that reads the web shouldn't also be able to email arbitrary addresses.
- Human approval for sensitive actions (Q11).
- Isolation / dual-LLM patterns: a "quarantined" model processes untrusted content and returns only structured, constrained data to the privileged agent. Research such as CaMeL formalises this with explicit data-flow control.
- Output and egress controls: allow-list destinations for emails and URLs; block rendering of untrusted links and images (a common exfiltration channel).
- Input scanning / classifiers: helpful, but bypassable. Treat them as one layer, not the solution.
- Marking untrusted content (delimiters, "this is data, not instructions") reduces risk but does not eliminate it.
- Monitoring: alert on unusual tool sequences.
Follow-ups to expect
- Walk through the "lethal trifecta." (See Q40.)
Related
Every expert started right here.