1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What is a Mixture of Experts (MoE) model?
30-second answerSay your answer out loud first, then reveal.

Example. Mixtral 8×7B has ~47B total parameters but ~13B active per token (2 of 8 experts per layer; the attention layers are shared). Several recent frontier and open models (e.g. DeepSeek-V3) use fine-grained MoE with many small experts plus shared experts.
Benefits
- More capacity per unit of compute: better quality at the same inference FLOPs.
- Faster training to a given loss for a fixed compute budget.
Challenges
- Memory: all experts must be in memory, so a 47B MoE needs memory like a 47B dense model even though it computes like 13B.
- Load balancing: the router may overuse a few experts. Fixed with auxiliary balancing losses or bias-based balancing.
- Serving: expert parallelism across GPUs; all-to-all communication; batch efficiency depends on routing.
- Fine-tuning can be less stable.
Interview nuance. MoE trades memory for compute. Great at large scale with high batch throughput; less attractive for single-GPU edge deployment.
Related
You understood something today that you didn't yesterday.