Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q30IntermediateConcept

What should you alert on for an LLM application, and how do you avoid alert fatigue?

30-second answerSay your answer out loud first, then reveal.

Alert design

AlertSeverityAction
Availability SLO burn rate high (1h window)PageCheck provider status, fail over, roll back recent deploy
p95 TTFT > SLO for 15 minPageInvestigate capacity, provider, recent changes
Parse/validation failure rate > 3x baselineTicket / page if critical flowCheck prompt or model changes
Groundedness pass rate down > 5 pts (daily)TicketError analysis; check index freshness
PII leak detector triggeredPage (security)Incident process
Daily cost > 150% forecastTicket + notify ownerLook for loops, abuse, config errors
Index sync lag > SLATicketFix connector

Avoiding fatigue

  • Multi-window burn-rate alerts instead of static thresholds.
  • Deduplicate and group related alerts.
  • Every page has an owner and a runbook link; review noisy alerts in retrospectives and fix or delete them.
  • Quality metrics (statistical, slower) feed tickets, not pages, unless the drop is severe.

Every expert started right here.