Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q22IntermediateConcept

How do you evaluate an AI agent?

30-second answerSay your answer out loud first, then reveal.
An eval pipeline: the eval dataset runs through the agent, the traces go to outcome checks, trajectory checks and an LLM judge, and the results form a scorecard per version that is checked for regression.

1. Outcome / end-to-end

  • Best is to check the environment state: was the ticket created with the right fields? Do the tests pass? Is the DB row updated?
  • Otherwise compare against a reference answer (exact match or a judge).

2. Trajectory

  • Did it call the required tools? Did it avoid forbidden ones?
  • Number of steps, tokens, cost and latency.
  • Policy adherence: did it ask for approval before refunding?
  • Don't over-specify the exact path. Agents can legitimately solve tasks in different ways.

3. Component / unit evals

  • Tool-selection accuracy, argument correctness.
  • Retrieval metrics (precision/recall).
  • Single-turn decisions at a fixed state.

Practical points that impress:

  • Non-determinism: run each task several times and report pass rates. The pass^k idea (passes all k trials) measures reliability, not just one lucky run.
  • Simulated users: for conversational agents, use an LLM playing the user with a persona and goal (as in benchmarks like τ-bench).
  • Start small: 20–50 real failure cases beat 1,000 synthetic ones. Grow the dataset from production failures.
  • Tools: LangSmith, Braintrust, DeepEval, Promptfoo, Ragas, or custom harnesses.

Follow-ups to expect

  • Pitfalls of LLM-as-judge? (See Q33.)

You understood something today that you didn't yesterday.