Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q33IntermediateScenario

Users say the assistant "got worse" over the last few weeks, but nobody changed anything. How do you investigate, and how do you detect this earlier next time?

30-second answerSay your answer out loud first, then reveal.

Investigation checklist

  1. Versions: were models pinned? Did the provider update the model behind "latest"? Any SDK, tokenizer or parser library upgrades?
  2. Data: index freshness, failed sync jobs, duplicated or conflicting documents, changed source formats breaking parsing.
  3. Inputs: compare query topic distributions (embedding clusters) between then and now. New product launches create questions the knowledge base can't answer.
  4. System: latency increases causing timeouts and truncated answers; more fallbacks to a weaker model; rate limits.
  5. Feedback data: which segments degraded (language, feature, tenant)?
  6. Reproduce: rerun the golden set now vs last month's results; diff the failing cases.

Prevention (continuous monitoring)

MonitorHow
Scheduled evalsNightly golden-set run against production config; alert on regression
Input driftTopic cluster distribution, unknown-intent rate, new-term frequency
Output driftLength, refusal rate, citation rate, format errors, sentiment
Retrieval healthTop-1 similarity score distribution, no-result rate
Sampled qualityLLM-judge on X% of traffic, plus weekly human review
Change logEvery model, prompt, data and config change recorded with timestamp

Slow is fine. Stopping is the only problem.