Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q37HardSystem design

Design online evaluation and quality monitoring for an LLM application in production.

30-second answerSay your answer out loud first, then reveal.
Production traces feed drift detectors, user signals and a sampler that sends traffic to async evaluators; their results land in a metrics store that drives alerts to on-call owners, while uncertain or severe cases go to a human review queue and on into offline eval datasets.

Design details

  1. Sampling: e.g. 2% random + 100% of thumbs-down + 100% of guardrail-flagged + oversampling of rare but important segments. Weight metrics back to the population.
  2. Evaluator costs: judges cost tokens. Budget them, use smaller judges for simple checks, and batch processing.
  3. Latency independence: evaluation is asynchronous, so it doesn't slow user requests.
  4. Version tagging: every metric is attributable to prompt, model and retrieval versions, for fast root-causing after deploys.
  5. Alerting: statistical thresholds (e.g. faithfulness failure rate exceeds its 7-day baseline + 3σ), not noisy absolute triggers.
  6. Privacy: PII scrubbing before evaluator calls if judges are external; retention limits.
  7. Feedback loop: human-confirmed failures become regression tests.

Dashboards: quality per criterion, safety incidents, user feedback rates, latency and cost, all per feature and version.

You understood something today that you didn't yesterday.