MoE (Mixture of Experts)
MoE (Mixture of Experts) is a "sparse activation for larger parameter count" model architecture: replace the FFN layer in Transformer with N "expert" sub-networks, each token activates only top-k experts (usually k=2).
Why MoE is needed
Dense models compute all parameters for every token generated. Larger model = more FLOPs = slower inference.
MoE makes total parameters large (e.g. 671B), but each token activates only a small portion (37B). Result:
- Training compute ≈ Dense 37B model (not 671B)
- Total capacity ≈ Dense 671B model (each expert learns different patterns)
- Inference FLOPs ≈ Dense 37B model (single token only goes through some experts)
Key components
- Expert networks: N independent FFNs (typically 8-256).
- Router/Gate: small linear layer + softmax, decides which experts each token goes to.
- Load balancing: training loss penalty to ensure experts are used evenly (avoid "celebrity experts").
Representative models
| Model | Total params | Active params | Experts | top-k |
|---|---|---|---|---|
| Mixtral 8x7B | 47B | 13B | 8 | 2 |
| DeepSeek V3 | 671B | 37B | 256 | 8 |
| GPT-4 (speculated) | ~1.8T | ~280B | ? | ? |
| Qwen-MoE | various | various | 60 | 4 |
Challenges
- Memory is the problem: all experts must live in GPU memory (MoE's "small compute, big memory" flips).
- Routing imbalance: early MoE models often collapsed to dense due to expert monopolies.
- Inference scheduling complexity: different tokens in a batch may go to different experts, dynamic scheduling required.
- Communication overhead: in distributed training, expert cross-GPU communication is a bottleneck.
When to use
- Training resources abundant, inference tight → MoE (large total, low per-token FLOPs).
- Inference tight but want fast inference → Dense (simple inference, batch-friendly).
- Edge deployment → almost never MoE (all experts must be loaded).