1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Estimate the GPUs needed to serve a 70B model to 32 concurrent users with 8K-token contexts. Walk through the math.
30-second answerSay your answer out loud first, then reveal.
Step 1 — Weights:
| Precision | Bytes/param | 70B weights |
|---|---|---|
| FP16/BF16 | 2 | 140 GB |
| FP8/INT8 | 1 | 70 GB |
| INT4 | ~0.5 | ~35–40 GB |
Step 2 — KV cache per token:
2 × layers × kv_heads × head_dim × bytes
= 2 × 80 × 8 × 128 × 2 (FP16) = 327,680 B ≈ 0.33 MB/tokenStep 3 — KV cache total:
32 users × 8,192 tokens × 0.33 MB ≈ 86 GB (FP16) | ≈ 43 GB (FP8 KV)(Worst case, assuming every user is at full context; real average usage is often lower, and paged allocation exploits that.)
Step 4 — Total + overhead:
| Config | Weights | KV | +~15% overhead | GPUs (80 GB) |
|---|---|---|---|---|
| FP16 / FP16 KV | 140 | 86 | ≈ 260 GB | 4 (TP=4) |
| FP8 / FP8 KV | 70 | 43 | ≈ 130 GB | 2 (TP=2) |
| INT4 / FP8 KV | ~38 | 43 | ≈ 93 GB | 2, or 1 large-memory GPU |
Step 5 — Throughput sanity check:
- Decode is memory-bandwidth bound. Each step reads the weights once for the whole batch.
- Rough single-GPU upper bound: tokens/step-time ≈ HBM bandwidth ÷ bytes read per step. Tensor parallelism splits the reads across GPUs.
- Confirm TTFT and TPOT under load with a benchmark (e.g. vLLM's benchmarking scripts) using a realistic prompt/output length distribution.
Trade-offs to discuss: quantization quality loss vs GPU cost; more GPUs for headroom (traffic spikes); a smaller model if quality allows; prefix caching if users share long system prompts.
Related
Little by little, you're building something great.