1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain data parallelism, tensor parallelism, pipeline parallelism, and ZeRO/FSDP.
30-second answerSay your answer out loud first, then reveal.
| Strategy | What's split | Communication | Best for | Limitation |
|---|---|---|---|---|
| Data parallel | Batch | All-reduce gradients each step | Scaling throughput | Each GPU must fit the full model + optimizer |
| ZeRO-1/2/3 / FSDP | Optimizer states (1), + gradients (2), + parameters (3) | More collective ops (all-gather params) | Training models that don't fit per GPU | Communication overhead |
| Tensor parallel | Matrices inside layers (e.g. split attention heads / MLP columns) | All-reduce per layer (frequent) | Huge layers; within a node | Needs NVLink-class bandwidth |
| Pipeline parallel | Layers into stages | Activations between stages | Very deep models, across nodes | "Bubble" idle time; micro-batch scheduling |
| Sequence / context parallel | Sequence length | Attention communication (ring) | Very long contexts | Complexity |
| Expert parallel | MoE experts across GPUs | All-to-all token routing | MoE models | Load balance |
Typical layout: TP within a node (8 GPUs over NVLink), PP across a few nodes, DP/FSDP across the rest.
Memory-saving extras
- Activation (gradient) checkpointing: recompute activations in the backward pass instead of storing them (trades compute for memory).
- Mixed precision (BF16 / FP8).
- CPU/NVMe offload (ZeRO-Offload) for limited hardware.
Why ZeRO matters: with plain DP, every GPU stores ~16 bytes per parameter (Q13). ZeRO-3 divides that by the number of GPUs, so a 70B model's ~1.1 TB of training state spreads across 64 GPUs as ~18 GB each (plus activations).
Related
This is what real progress feels like.