1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your path
You're running the launch-readiness review for a new customer-facing AI feature. What evaluation and guardrail evidence do you require before go-live?
30-second answerSay your answer out loud first, then reveal.
Launch checklist (example thresholds)
| Area | Requirement |
|---|---|
| Task quality | ≥ 90% pass on 500-case golden set (95% CI lower bound ≥ 87%); no slice < 85% |
| Critical errors | < 0.5% critical errors (wrong amounts, policy violations) |
| Safety | Red-team suite: 0 open critical findings; harmful-pass rate < 1% on adversarial set |
| Over-refusal | False refusal < 3% on benign sensitive set |
| Judges | Calibrated: TNR ≥ 0.85 vs human labels |
| Guardrails | PII leak tests 100% pass; added latency p95 < 300ms |
| Performance | p95 end-to-end latency within SLA under 2x expected load |
| Cost | Cost per conversation within budget at forecast volume |
| Privacy / security | DPIA done; security review sign-off; data retention configured |
| Operations | Dashboards + alerts; on-call runbook; rollback and kill switch tested |
| Human fallback | Escalation to human agents works; SLA defined |
| UX | AI disclosure, feedback buttons, clear limitations messaging |
| Rollout plan | 1% → 10% → 50% → 100%, with go/no-go metrics and an owner at each step |
Process: a single owner compiles the evidence; reviewers from product, engineering, security, legal/compliance and support sign off; exceptions are documented with risk acceptance by an accountable leader.
Interview signal. Concrete thresholds, cross-functional sign-off, and staged rollout with stop criteria. That turns "it seems ready" into a defensible decision.
Related
Every expert started right here.