Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q36HardScenario

Estimate the GPUs needed to serve a 70B model to 32 concurrent users with 8K-token contexts. Walk through the math.

30-second answerSay your answer out loud first, then reveal.

Step 1 — Weights:

PrecisionBytes/param70B weights
FP16/BF162140 GB
FP8/INT8170 GB
INT4~0.5~35–40 GB

Step 2 — KV cache per token:

text
2 × layers × kv_heads × head_dim × bytes
= 2 × 80 × 8 × 128 × 2 (FP16) = 327,680 B ≈ 0.33 MB/token

Step 3 — KV cache total:

text
32 users × 8,192 tokens × 0.33 MB ≈ 86 GB (FP16)  |  ≈ 43 GB (FP8 KV)

(Worst case, assuming every user is at full context; real average usage is often lower, and paged allocation exploits that.)

Step 4 — Total + overhead:

ConfigWeightsKV+~15% overheadGPUs (80 GB)
FP16 / FP16 KV14086≈ 260 GB4 (TP=4)
FP8 / FP8 KV7043≈ 130 GB2 (TP=2)
INT4 / FP8 KV~3843≈ 93 GB2, or 1 large-memory GPU

Step 5 — Throughput sanity check:

  • Decode is memory-bandwidth bound. Each step reads the weights once for the whole batch.
  • Rough single-GPU upper bound: tokens/step-time ≈ HBM bandwidth ÷ bytes read per step. Tensor parallelism splits the reads across GPUs.
  • Confirm TTFT and TPOT under load with a benchmark (e.g. vLLM's benchmarking scripts) using a realistic prompt/output length distribution.

Trade-offs to discuss: quantization quality loss vs GPU cost; more GPUs for headroom (traffic spikes); a smaller model if quality allows; prefix caching if users share long system prompts.

Little by little, you're building something great.