1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Users say the assistant "got worse" over the last few weeks, but nobody changed anything. How do you investigate, and how do you detect this earlier next time?
30-second answerSay your answer out loud first, then reveal.
Investigation checklist
- Versions: were models pinned? Did the provider update the model behind "latest"? Any SDK, tokenizer or parser library upgrades?
- Data: index freshness, failed sync jobs, duplicated or conflicting documents, changed source formats breaking parsing.
- Inputs: compare query topic distributions (embedding clusters) between then and now. New product launches create questions the knowledge base can't answer.
- System: latency increases causing timeouts and truncated answers; more fallbacks to a weaker model; rate limits.
- Feedback data: which segments degraded (language, feature, tenant)?
- Reproduce: rerun the golden set now vs last month's results; diff the failing cases.
Prevention (continuous monitoring)
| Monitor | How |
|---|---|
| Scheduled evals | Nightly golden-set run against production config; alert on regression |
| Input drift | Topic cluster distribution, unknown-intent rate, new-term frequency |
| Output drift | Length, refusal rate, citation rate, format errors, sentiment |
| Retrieval health | Top-1 similarity score distribution, no-result rate |
| Sampled quality | LLM-judge on X% of traffic, plus weekly human review |
| Change log | Every model, prompt, data and config change recorded with timestamp |
Related
Slow is fine. Stopping is the only problem.