Skip to main content
ZICQ

Wiki Concepts

MoE (Mixture of Experts)

Concepts
Aliases: Mixture of Experts ·2026-09-14

MoE (Mixture of Experts)

MoE (Mixture of Experts) is a "sparse activation for larger parameter count" model architecture: replace the FFN layer in Transformer with N "expert" sub-networks, each token activates only top-k experts (usually k=2).

Why MoE is needed

Dense models compute all parameters for every token generated. Larger model = more FLOPs = slower inference.

MoE makes total parameters large (e.g. 671B), but each token activates only a small portion (37B). Result:

  • Training compute ≈ Dense 37B model (not 671B)
  • Total capacity ≈ Dense 671B model (each expert learns different patterns)
  • Inference FLOPs ≈ Dense 37B model (single token only goes through some experts)

Key components

  • Expert networks: N independent FFNs (typically 8-256).
  • Router/Gate: small linear layer + softmax, decides which experts each token goes to.
  • Load balancing: training loss penalty to ensure experts are used evenly (avoid "celebrity experts").

Representative models

Model Total params Active params Experts top-k
Mixtral 8x7B 47B 13B 8 2
DeepSeek V3 671B 37B 256 8
GPT-4 (speculated) ~1.8T ~280B ? ?
Qwen-MoE various various 60 4

Challenges

  • Memory is the problem: all experts must live in GPU memory (MoE's "small compute, big memory" flips).
  • Routing imbalance: early MoE models often collapsed to dense due to expert monopolies.
  • Inference scheduling complexity: different tokens in a batch may go to different experts, dynamic scheduling required.
  • Communication overhead: in distributed training, expert cross-GPU communication is a bottleneck.

When to use

  • Training resources abundant, inference tight → MoE (large total, low per-token FLOPs).
  • Inference tight but want fast inference → Dense (simple inference, batch-friendly).
  • Edge deployment → almost never MoE (all experts must be loaded).