1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you load test an LLM service realistically?
30-second answerSay your answer out loud first, then reveal.
Steps
- Build a workload profile from production logs: distributions of prompt tokens, output tokens, features and request rate by hour.
- Generate traffic with replayed (anonymised) or synthetic prompts matching those distributions. Identical prompts overstate cache benefits.
- Ramp load in steps; record TTFT, ITL and end-to-end p50/p95/p99, error rates, throughput, GPU and KV-cache utilisation.
- Find the knee: the max load at which SLOs still hold, which is your per-replica capacity for planning.
- Stress and chaos: spikes (3x in 1 minute), a replica kill, provider 429 injection, slow vector DB.
- Soak test: hours at steady load to find memory leaks and gradual degradation.
Tools: vLLM/engine benchmark scripts, Locust or k6 with streaming support, provider-specific load-test guidance (respect rate limits and terms; use test quotas).
Common mistakes
- Load testing with short identical prompts, which gives unrealistically good latency and cache hit rates.
Related
Slow is fine. Stopping is the only problem.