Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q48HardSystem design

Design a human review and annotation operation that supports evals at scale (thousands of items per week).

30-second answerSay your answer out loud first, then reveal.
Items are prioritised and sampled, mixed with gold items into task queues by skill, reviewed in a review UI, checked by quality control, and turned into approved labels for eval datasets and judge calibration plus reviewer metrics that drive training and guideline updates.

Key elements

  1. Guidelines: definitions, decision trees, many examples (especially borderline ones), versioned; updated when disagreements reveal ambiguity.
  2. Reviewer management: qualification tests, ongoing gold-question accuracy, calibration sessions, wellbeing support (for reviewers exposed to harmful content).
  3. Throughput and cost: estimate minutes per item × volume; for complex items, use LLM pre-labelling with human verification (faster, but watch anchoring bias).
  4. Agreement metrics: Cohen's / Krippendorff's alpha; disagreement queues to senior adjudicators.
  5. Privacy: PII redaction where possible, access controls, secure environments, NDAs.
  6. Closing the loop: label distributions and new failure categories feed back into error analysis and product teams.

Slow is fine. Stopping is the only problem.