1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What should you alert on for an LLM application, and how do you avoid alert fatigue?
30-second answerSay your answer out loud first, then reveal.
Alert design
| Alert | Severity | Action |
|---|---|---|
| Availability SLO burn rate high (1h window) | Page | Check provider status, fail over, roll back recent deploy |
| p95 TTFT > SLO for 15 min | Page | Investigate capacity, provider, recent changes |
| Parse/validation failure rate > 3x baseline | Ticket / page if critical flow | Check prompt or model changes |
| Groundedness pass rate down > 5 pts (daily) | Ticket | Error analysis; check index freshness |
| PII leak detector triggered | Page (security) | Incident process |
| Daily cost > 150% forecast | Ticket + notify owner | Look for loops, abuse, config errors |
| Index sync lag > SLA | Ticket | Fix connector |
Avoiding fatigue
- Multi-window burn-rate alerts instead of static thresholds.
- Deduplicate and group related alerts.
- Every page has an owner and a runbook link; review noisy alerts in retrospectives and fix or delete them.
- Quality metrics (statistical, slower) feed tickets, not pages, unless the drop is severe.
Related
Every expert started right here.