1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What metrics do you use to evaluate an AI agent's tool use and trajectory?
30-second answerSay your answer out loud first, then reveal.
Metric set
| Metric | Definition |
|---|---|
| Task success rate | Final state / answer correct (state checks preferred) |
| pass^k | Succeeds in all k trials (consistency) |
| Tool-selection accuracy | Correct tool chosen at decision points (from labelled steps) |
| Argument accuracy | Tool args valid and correct (schema + semantic checks) |
| Required actions present | E.g. "verify identity before refund" |
| Forbidden actions absent | E.g. no refund above limit, no external email |
| Efficiency | Steps, tool calls, tokens, cost, wall-clock time |
| Recovery rate | Succeeds despite injected tool errors |
| Escalation correctness | Hands off to a human when it should |
Implementation: sandboxed environment with seeded data; mocked or real tools with fault injection; a trace capture for each run; a scorecard per agent version. (See the Agentic AI guide for design patterns.)
Common mistakes
- Exact trajectory matching (penalises valid alternative paths).
- Only judging the final message, while missing a harmful action taken midway.
Related
You understood something today that you didn't yesterday.