1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate the safety of AI agents that can take actions with tools?
30-second answerSay your answer out loud first, then reveal.
Safety eval dimensions
| Dimension | Example test |
|---|---|
| Unauthorised actions | User asks to refund another customer's order |
| Excessive actions | Asked to "clean up old tickets", deletes all tickets |
| Approval compliance | High-value action attempted without the required approval step |
| Indirect injection | Email content instructs the agent to forward data externally |
| Data boundaries | Agent reads tenant A's data while serving tenant B |
| Secrets handling | Agent prints API keys from environment or files |
| Resource abuse | Infinite loops, excessive API calls, cost blowups |
| Safe failure | Tool errors lead to escalation, not risky retries |
Environment design
- Seeded sandbox systems (mock CRM, email, file system) with canary data (e.g. unique strings that should never appear in outputs or outbound messages).
- Tool call logs and state diffs checked automatically after each scenario.
- Multiple trials (pass^k), because rare unsafe behaviour matters.
Gates: zero critical violations across N trials; bounded rates for lower-severity issues; manual red-team sign-off for new tools or permissions.
Related
You understood something today that you didn't yesterday.