Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q28IntermediateConcept

How do you load test an LLM service realistically?

30-second answerSay your answer out loud first, then reveal.

Steps

  1. Build a workload profile from production logs: distributions of prompt tokens, output tokens, features and request rate by hour.
  2. Generate traffic with replayed (anonymised) or synthetic prompts matching those distributions. Identical prompts overstate cache benefits.
  3. Ramp load in steps; record TTFT, ITL and end-to-end p50/p95/p99, error rates, throughput, GPU and KV-cache utilisation.
  4. Find the knee: the max load at which SLOs still hold, which is your per-replica capacity for planning.
  5. Stress and chaos: spikes (3x in 1 minute), a replica kill, provider 429 injection, slow vector DB.
  6. Soak test: hours at steady load to find memory leaks and gradual degradation.

Tools: vLLM/engine benchmark scripts, Locust or k6 with streaming support, provider-specific load-test guidance (respect rate limits and terms; use test quotas).

Common mistakes

  • Load testing with short identical prompts, which gives unrealistically good latency and cache hit rates.

Slow is fine. Stopping is the only problem.