1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Design an evaluation harness and CI/CD process for an agent.
30-second answerSay your answer out loud first, then reveal.

Components
- Versioned config: the agent version = prompt + tools + model + parameters, all in git.
- Datasets (curated, growing):
• Golden tasks: realistic tasks with verifiable outcomes.
• Regression set: every fixed production bug.
• Edge cases: ambiguity, missing data, tool failures (injected faults).
• Safety: injections, policy violations, PII handling. - Environment: sandboxed replicas or mocks of tools; seeded test databases so outcome checks are deterministic; recorded responses for external APIs.
- Graders: state checks (preferred), unit tests, schema checks, trajectory assertions, calibrated LLM judges (Q33).
- Statistics: multiple trials per task, because of non-determinism. Report confidence intervals and don't overreact to ±1% noise.
- Gates: e.g. success ≥ baseline − 1%, zero critical safety failures, cost ≤ +10%.
- Reports: per-task diffs showing which tasks flipped from pass to fail, with trace links.
Cost tip: run the fast tier on every PR and the expensive full tier nightly or before release.
Related
This is what real progress feels like.