Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q22IntermediateConcept

What metrics do you use to evaluate an AI agent's tool use and trajectory?

30-second answerSay your answer out loud first, then reveal.

Metric set

MetricDefinition
Task success rateFinal state / answer correct (state checks preferred)
pass^kSucceeds in all k trials (consistency)
Tool-selection accuracyCorrect tool chosen at decision points (from labelled steps)
Argument accuracyTool args valid and correct (schema + semantic checks)
Required actions presentE.g. "verify identity before refund"
Forbidden actions absentE.g. no refund above limit, no external email
EfficiencySteps, tool calls, tokens, cost, wall-clock time
Recovery rateSucceeds despite injected tool errors
Escalation correctnessHands off to a human when it should

Implementation: sandboxed environment with seeded data; mocked or real tools with fault injection; a trace capture for each run; a scorecard per agent version. (See the Agentic AI guide for design patterns.)

Common mistakes

  • Exact trajectory matching (penalises valid alternative paths).
  • Only judging the final message, while missing a harmful action taken midway.

You understood something today that you didn't yesterday.