Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q27IntermediateConcept

What is a Mixture of Experts (MoE) model?

30-second answerSay your answer out loud first, then reveal.
A token's hidden state goes to a router that picks the top-2 experts (Expert 2 and Expert 5) while the others (Expert 1, Expert 8) are not used; the chosen experts' outputs are combined in a weighted sum by router scores to give the output.
Example. Mixtral 8×7B has ~47B total parameters but ~13B active per token (2 of 8 experts per layer; the attention layers are shared). Several recent frontier and open models (e.g. DeepSeek-V3) use fine-grained MoE with many small experts plus shared experts.

Benefits

  • More capacity per unit of compute: better quality at the same inference FLOPs.
  • Faster training to a given loss for a fixed compute budget.

Challenges

  1. Memory: all experts must be in memory, so a 47B MoE needs memory like a 47B dense model even though it computes like 13B.
  2. Load balancing: the router may overuse a few experts. Fixed with auxiliary balancing losses or bias-based balancing.
  3. Serving: expert parallelism across GPUs; all-to-all communication; batch efficiency depends on routing.
  4. Fine-tuning can be less stable.
Interview nuance. MoE trades memory for compute. Great at large scale with high batch throughput; less attractive for single-GPU edge deployment.

You understood something today that you didn't yesterday.