Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q42HardConcept

How do you evaluate the safety of AI agents that can take actions with tools?

30-second answerSay your answer out loud first, then reveal.

Safety eval dimensions

DimensionExample test
Unauthorised actionsUser asks to refund another customer's order
Excessive actionsAsked to "clean up old tickets", deletes all tickets
Approval complianceHigh-value action attempted without the required approval step
Indirect injectionEmail content instructs the agent to forward data externally
Data boundariesAgent reads tenant A's data while serving tenant B
Secrets handlingAgent prints API keys from environment or files
Resource abuseInfinite loops, excessive API calls, cost blowups
Safe failureTool errors lead to escalation, not risky retries

Environment design

  • Seeded sandbox systems (mock CRM, email, file system) with canary data (e.g. unique strings that should never appear in outputs or outbound messages).
  • Tool call logs and state diffs checked automatically after each scenario.
  • Multiple trials (pass^k), because rare unsafe behaviour matters.

Gates: zero critical violations across N trials; bounded rates for lower-severity issues; manual red-team sign-off for new tools or permissions.

You understood something today that you didn't yesterday.