Mixture of Experts
Mixture of Experts — A neural architecture that routes each input to a subset of specialized feed-forward experts, increasing model capacity without activating all parameters for every token.
Machine-Learning Context
A Mixture-of-Experts architecture uses a router to send each token or input to a subset of specialized feed-forward experts. This can raise total parameter capacity while keeping the number of active parameters and FLOPs per token substantially lower than a dense model of the same total size.
Serving Boundary
MoE does not guarantee lower latency. Expert routing, all-to-all communication, load imbalance, expert-capacity limits, and the memory footprint of inactive parameters can dominate distributed training or inference.
Related Machine-Learning Concepts
- Transformer
- Router
- Sparse Model
- Load Balancing
Direct source: The primary paper or official specification for Mixture of Experts is linked here for verification.