Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your path

Q50HardScenario

You're running the launch-readiness review for a new customer-facing AI feature. What evaluation and guardrail evidence do you require before go-live?

30-second answerSay your answer out loud first, then reveal.

Launch checklist (example thresholds)

AreaRequirement
Task quality≥ 90% pass on 500-case golden set (95% CI lower bound ≥ 87%); no slice < 85%
Critical errors< 0.5% critical errors (wrong amounts, policy violations)
SafetyRed-team suite: 0 open critical findings; harmful-pass rate < 1% on adversarial set
Over-refusalFalse refusal < 3% on benign sensitive set
JudgesCalibrated: TNR ≥ 0.85 vs human labels
GuardrailsPII leak tests 100% pass; added latency p95 < 300ms
Performancep95 end-to-end latency within SLA under 2x expected load
CostCost per conversation within budget at forecast volume
Privacy / securityDPIA done; security review sign-off; data retention configured
OperationsDashboards + alerts; on-call runbook; rollback and kill switch tested
Human fallbackEscalation to human agents works; SLA defined
UXAI disclosure, feedback buttons, clear limitations messaging
Rollout plan1% → 10% → 50% → 100%, with go/no-go metrics and an owner at each step

Process: a single owner compiles the evidence; reviewers from product, engineering, security, legal/compliance and support sign off; exceptions are documented with risk acceptance by an accountable leader.

Interview signal. Concrete thresholds, cross-functional sign-off, and staged rollout with stop criteria. That turns "it seems ready" into a defensible decision.

Every expert started right here.