Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q45HardScenario

Your self-hosted model servers degrade gradually: latency creeps up over days until restarts fix it. How do you debug it?

30-second answerSay your answer out loud first, then reveal.

Investigation

  1. Metrics vs uptime: plot latency, CPU RSS, GPU memory, KV-cache utilisation, open connections and file descriptors per replica against time since restart. Which grows?
  2. Compare replicas: do all degrade, or only some (a hardware issue, e.g. GPU thermal throttling or ECC errors in DCGM metrics)?
  3. Traffic correlation: does degradation accelerate with certain request types (long contexts, specific adapters, image inputs)?
  4. Reproduce: a soak test in staging with production-like traffic; bisect engine versions or config flags.
  5. Profile: Python memory profilers for CPU-side leaks; engine debug logs; GPU memory snapshots.
  6. Check upstream: release notes and issues for the inference engine, CUDA and driver versions.

Typical fixes

CauseFix
Engine bug (known memory leak)Upgrade or downgrade engine version
Unbounded cachesConfigure cache size limits / eviction
Client connection leaksFix HTTP client pooling; timeouts
Thermal throttling / failing GPUReplace node; alert on DCGM health metrics
FragmentationEngine settings; periodic graceful restarts (temporary)

Mitigation while fixing: rolling restarts with connection draining, staggered across replicas; alerts on the leading indicator (memory growth rate).

Every expert started right here.