What Is a Mixture-of-Experts Model?
Mixture-of-experts is a neural-network design that splits a model into many specialised sub-networks and activates only a few of them per input, chosen by a learned router. It lets total parameter count — and capacity — grow without a proportional rise in compute per token, because most experts stay idle on any given input. Many frontier models use it to be large and affordable at once; the trade-offs are routing complexity and uneven expert utilisation.
Sparse activation
A dense model runs every parameter for every token. A mixture-of-experts routes each token to a small subset of experts, so a model with very high total capacity spends compute like a much smaller one. The router is trained alongside the experts to send each input where it will be handled best.
For the agent economy this matters indirectly but deeply: it is part of why capable models became cheap enough to run an agent workforce at all.
Related standards
Questions
Does mixture-of-experts make a model smarter?
Not by itself — it makes a model larger-capacity for the same inference cost, which can translate to better quality when trained well.
What is the hard part?
Routing: keeping experts balanced so some are not overused and others never trained, and avoiding instability during training.
Keep reading
By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.