What is a mixture-of-experts (MoE) model?
A mixture-of-experts (MoE) model is a neural network built from many parallel sub-networks, called experts, plus a small routing network that picks a handful of them to process each token. Unlike a dense model, which uses every one of its parameters on every token, an MoE model stores a large total parameter count but only activates a small fraction of it per token, which is why the architecture matters so much for how fast a model runs.
Active parameters vs total parameters
The distinction that actually matters when running an MoE model is active parameters versus total parameters. Total is the full set of weights stored on disk and in memory, the number in the model's name. Active is how many of those get read and computed on a given token, decided by the router. A model might be labeled "120B" as its total count while activating only 12B parameters per token, and that smaller active number, not the headline figure, is what predicts decode speed. GPUwerk's glossary covers this split in more depth, including measured examples, in its active vs total parameters section.
Why route tokens through experts at all
Training one enormous dense model gets expensive fast, since every added parameter costs compute on every single token, during training and at inference. Splitting the network into experts and routing each token to only a few of them decouples total capacity from per-token compute: the model can hold far more total parameters, spread across specialized experts, without paying the full compute cost of all of them on every token. The trade-off is added complexity, since the router has to learn which experts suit which kinds of tokens, and uneven expert usage can hurt training efficiency if not managed carefully.
What this means for hardware
Memory capacity still has to cover the full total parameter count, since every expert is loaded even though most sit idle on any given token. But the compute and memory-bandwidth cost per token tracks the active count, not the total. That combination, large memory footprint paired with comparatively light per-token compute, tends to suit unified-memory hardware well, since it favors a single large, fast memory pool over splitting the model across several GPUs' separate memory. See the glossary's MoE vs dense comparison for the underlying mechanics and links to GPUwerk's own measured throughput numbers.
MoE isn't the same thing as an agent
Mixture-of-experts describes a model's internal architecture, how it's built and how it computes an output. It's unrelated to whether a system is set up as an "agent," which is about how a model is used, with tools, multi-step reasoning, and autonomy layered on top. An MoE model can back a simple chat completion or an agentic system just as easily as a dense model can; see what is an LLM agent for that distinction.
"Experts" aren't specialized the way the name suggests
The term "expert" implies each sub-network handles a distinct topic, one for code, one for math, and so on, but that's rarely how it works out in practice. Routing tends to split along statistical patterns in the training data rather than clean human categories, so a given expert might end up specializing in a syntactic pattern or token type rather than a subject domain. The architecture is a compute-efficiency mechanism first; any topical specialization that emerges is a side effect of training, not a designed property.
Examples in current models
MoE architectures now appear across several open model families at a range of sizes, from tens of billions of total parameters up to several hundred billion, with active counts typically a fraction of the total. The specific ratio of active to total parameters varies by model and is usually stated on the model's own card or release notes rather than being a fixed convention across the category, so it's worth checking the specific model rather than assuming a ratio.