Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q44HardSystem design

Design an evaluation harness and CI/CD process for an agent.

30-second answerSay your answer out loud first, then reveal.
An agent CI/CD flow: a PR runs fast evals, then on merge a full sandboxed suite produces a scorecard; if the gates are met it goes to a canary deploy and production monitoring, otherwise it is blocked with a diff report, and failed production traces are labelled and added to the eval datasets that feed both eval tiers.

Components

  1. Versioned config: the agent version = prompt + tools + model + parameters, all in git.
  2. Datasets (curated, growing):
    • Golden tasks: realistic tasks with verifiable outcomes.
    • Regression set: every fixed production bug.
    • Edge cases: ambiguity, missing data, tool failures (injected faults).
    • Safety: injections, policy violations, PII handling.
  3. Environment: sandboxed replicas or mocks of tools; seeded test databases so outcome checks are deterministic; recorded responses for external APIs.
  4. Graders: state checks (preferred), unit tests, schema checks, trajectory assertions, calibrated LLM judges (Q33).
  5. Statistics: multiple trials per task, because of non-determinism. Report confidence intervals and don't overreact to ±1% noise.
  6. Gates: e.g. success ≥ baseline − 1%, zero critical safety failures, cost ≤ +10%.
  7. Reports: per-task diffs showing which tasks flipped from pass to fail, with trace links.

Cost tip: run the fast tier on every PR and the expensive full tier nightly or before release.

This is what real progress feels like.