1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your self-hosted model servers degrade gradually: latency creeps up over days until restarts fix it. How do you debug it?
30-second answerSay your answer out loud first, then reveal.
Investigation
- Metrics vs uptime: plot latency, CPU RSS, GPU memory, KV-cache utilisation, open connections and file descriptors per replica against time since restart. Which grows?
- Compare replicas: do all degrade, or only some (a hardware issue, e.g. GPU thermal throttling or ECC errors in DCGM metrics)?
- Traffic correlation: does degradation accelerate with certain request types (long contexts, specific adapters, image inputs)?
- Reproduce: a soak test in staging with production-like traffic; bisect engine versions or config flags.
- Profile: Python memory profilers for CPU-side leaks; engine debug logs; GPU memory snapshots.
- Check upstream: release notes and issues for the inference engine, CUDA and driver versions.
Typical fixes
| Cause | Fix |
|---|---|
| Engine bug (known memory leak) | Upgrade or downgrade engine version |
| Unbounded caches | Configure cache size limits / eviction |
| Client connection leaks | Fix HTTP client pooling; timeouts |
| Thermal throttling / failing GPU | Replace node; alert on DCGM health metrics |
| Fragmentation | Engine settings; periodic graceful restarts (temporary) |
Mitigation while fixing: rolling restarts with connection draining, staggered across replicas; alerts on the leading indicator (memory growth rate).
Related
Every expert started right here.