1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate an AI agent?
30-second answerSay your answer out loud first, then reveal.

1. Outcome / end-to-end
- Best is to check the environment state: was the ticket created with the right fields? Do the tests pass? Is the DB row updated?
- Otherwise compare against a reference answer (exact match or a judge).
2. Trajectory
- Did it call the required tools? Did it avoid forbidden ones?
- Number of steps, tokens, cost and latency.
- Policy adherence: did it ask for approval before refunding?
- Don't over-specify the exact path. Agents can legitimately solve tasks in different ways.
3. Component / unit evals
- Tool-selection accuracy, argument correctness.
- Retrieval metrics (precision/recall).
- Single-turn decisions at a fixed state.
Practical points that impress:
- Non-determinism: run each task several times and report pass rates. The pass^k idea (passes all k trials) measures reliability, not just one lucky run.
- Simulated users: for conversational agents, use an LLM playing the user with a persona and goal (as in benchmarks like τ-bench).
- Start small: 20–50 real failure cases beat 1,000 synthetic ones. Grow the dataset from production failures.
- Tools: LangSmith, Braintrust, DeepEval, Promptfoo, Ragas, or custom harnesses.
Follow-ups to expect
- Pitfalls of LLM-as-judge? (See Q33.)
Related
You understood something today that you didn't yesterday.