Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q39HardConcept

Explain data parallelism, tensor parallelism, pipeline parallelism, and ZeRO/FSDP.

30-second answerSay your answer out loud first, then reveal.
StrategyWhat's splitCommunicationBest forLimitation
Data parallelBatchAll-reduce gradients each stepScaling throughputEach GPU must fit the full model + optimizer
ZeRO-1/2/3 / FSDPOptimizer states (1), + gradients (2), + parameters (3)More collective ops (all-gather params)Training models that don't fit per GPUCommunication overhead
Tensor parallelMatrices inside layers (e.g. split attention heads / MLP columns)All-reduce per layer (frequent)Huge layers; within a nodeNeeds NVLink-class bandwidth
Pipeline parallelLayers into stagesActivations between stagesVery deep models, across nodes"Bubble" idle time; micro-batch scheduling
Sequence / context parallelSequence lengthAttention communication (ring)Very long contextsComplexity
Expert parallelMoE experts across GPUsAll-to-all token routingMoE modelsLoad balance

Typical layout: TP within a node (8 GPUs over NVLink), PP across a few nodes, DP/FSDP across the rest.

Memory-saving extras

  • Activation (gradient) checkpointing: recompute activations in the backward pass instead of storing them (trades compute for memory).
  • Mixed precision (BF16 / FP8).
  • CPU/NVMe offload (ZeRO-Offload) for limited hardware.

Why ZeRO matters: with plain DP, every GPU stores ~16 bytes per parameter (Q13). ZeRO-3 divides that by the number of GPUs, so a 70B model's ~1.1 TB of training state spreads across 64 GPUs as ~18 GB each (plus activations).

This is what real progress feels like.