Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q36HardSystem design

Design an evaluation platform for an organisation with 20 LLM-powered features built by different teams.

30-second answerSay your answer out loud first, then reveal.
Production traces, synthetic generators and SME labels are curated into versioned datasets, which with the grader registry and per-team eval configs feed an execution engine whose results store drives the comparison UI, CI gates and alerts, and launch reviews.

Key design points

  1. Evals as code: each feature's eval config (dataset version, graders, thresholds) lives in its repo and is reviewed like code.
  2. Shared assets: org-wide safety and red-team sets, PII leak tests, and multilingual sets every feature must pass.
  3. Grader registry with calibration metadata: each judge records its agreement with humans (TPR/TNR) and the model version. Uncalibrated judges are flagged.
  4. Run efficiency: cache outputs per (input, system version); parallelise; run cheap subsets per PR and full suites nightly.
  5. Statistics built in: confidence intervals, paired comparisons, flipped-example views.
  6. Human review integration: low-agreement or high-stakes items go to labelling queues; labels flow back to the datasets.
  7. Governance: launch criteria templates (quality, safety, latency, cost), required for production releases.
  8. Access and privacy: datasets derived from production traces are PII-scrubbed, access-controlled and retained according to policy.

Success metrics: % of features with CI eval gates, regressions caught before release, time to set up evals for a new feature.

Little by little, you're building something great.