1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Design online evaluation and quality monitoring for an LLM application in production.
30-second answerSay your answer out loud first, then reveal.

Design details
- Sampling: e.g. 2% random + 100% of thumbs-down + 100% of guardrail-flagged + oversampling of rare but important segments. Weight metrics back to the population.
- Evaluator costs: judges cost tokens. Budget them, use smaller judges for simple checks, and batch processing.
- Latency independence: evaluation is asynchronous, so it doesn't slow user requests.
- Version tagging: every metric is attributable to prompt, model and retrieval versions, for fast root-causing after deploys.
- Alerting: statistical thresholds (e.g. faithfulness failure rate exceeds its 7-day baseline + 3σ), not noisy absolute triggers.
- Privacy: PII scrubbing before evaluator calls if judges are external; retention limits.
- Feedback loop: human-confirmed failures become regression tests.
Dashboards: quality per criterion, safety incidents, user feedback rates, latency and cost, all per feature and version.
Related
You understood something today that you didn't yesterday.