What Is a Mixture-of-Experts Model?

Mixture-of-experts is a neural-network design that splits a model into many specialised sub-networks and activates only a few of them per input, chosen by a learned router. It lets total parameter count — and capacity — grow without a proportional rise in compute per token, because most experts stay idle on any given input. Many frontier models use it to be large and affordable at once; the trade-offs are routing complexity and uneven expert utilisation.

Idea
Many experts, only a few active per input
Benefit
Capacity without proportional compute
Early work
Shazeer et al., 2017

Sparse activation

A dense model runs every parameter for every token. A mixture-of-experts routes each token to a small subset of experts, so a model with very high total capacity spends compute like a much smaller one. The router is trained alongside the experts to send each input where it will be handled best.

For the agent economy this matters indirectly but deeply: it is part of why capable models became cheap enough to run an agent workforce at all.

Related standards

Shazeer et al., 2017 — Sparsely-Gated Mixture-of-Experts

Questions

Does mixture-of-experts make a model smarter?

Not by itself — it makes a model larger-capacity for the same inference cost, which can translate to better quality when trained well.

What is the hard part?

Routing: keeping experts balanced so some are not overused and others never trained, and avoiding instability during training.

Keep reading

related
What Is a Vector Database?
related
Managing an Agent’s Context Window
related
What Is Agent Context (and the Context Window)?
related
What Is Prompt Injection?
referenced by
What Is an Embedding?
referenced by
What Is Fine-Tuning?
referenced by
What Is a Transformer (Neural Network Architecture)?
referenced by
What Is Model Distillation?
referenced by
What Is Model Quantization?
referenced by
Transformer vs RNN: What Changed?
Agent Data & Memory
What Is Agent Memory?
Agent Data & Memory
Retrieval-Augmented Generation (RAG) for Agents
Agent Data & Memory
What Is an Agent Knowledge Graph?
Agent Data & Memory
What Is Content Addressing?
Agent Data & Memory
What Is an AI Hallucination?
Agent Data & Memory
RAG vs Fine-Tuning: Which Should You Use?
Agent Data & Memory
What Is the Attention Mechanism?
Agent Data & Memory
What Is Chain-of-Thought Prompting?
Agent Data & Memory
What Is RLHF (Reinforcement Learning from Human Feedback)?
Agent Data & Memory
What Is Tokenization in AI (Subword Tokens)?
Agent Data & Memory
What Is Semantic Search?
Agent Data & Memory
RAG vs Long Context: How Should a Model Get Its Knowledge?
Agent Data & Memory
Supervised vs Unsupervised Learning: What’s the Difference?

By Michael Gord · published 2026-10-09 · part of the Agentic Encyclopedia. Dates are the day of publication; events are cited at their own dates.