A mixture-of-experts model is priced like a small model and provisioned like a huge one, because a router sends each token through only a few percent of the weights.
Why you'd careThe thing you have already noticed
You noticed that a model advertised at several hundred billion parameters answered faster than a 70B dense one, and also that nobody could run it on fewer than eight GPUs. Or you benchmarked one and got excellent single-request latency alongside disappointing throughput under mixed load. Mixture-of-experts explains both. The feed-forward block, roughly two thirds of a dense transformer's parameters, is replaced by many parallel copies plus a small router that picks a few of them per token. Mixtral 8x7B holds 46.7B parameters and runs about 12.9B of them for each token. DeepSeek-V3 holds 671B and runs about 37B. Compute follows the active count, memory follows the total, and every weight must be resident even though almost none will be touched by your token.
In and outWhat goes in, what comes out
| In | The normalised residual vector for one token, shape [d_model]. Routing is per token and per layer, so the same request routes differently at every MoE layer and two adjacent tokens routinely take different paths. |
|---|---|
| Process | A linear router produces one score per expert. Take the top k, normalise their gates, dispatch the token to those experts — an all-to-all network exchange when experts are sharded across GPUs — run each expert's feed-forward network, and combine outputs weighted by gate. |
| Out | A single [d_model] vector added back into the residual stream, indistinguishable in shape from a dense feed-forward output. Which experts fired is not encoded anywhere in it. |
The routing decision is discrete and then disappears. Downstream layers see a vector, not a path, so unless the server explicitly logs router decisions you cannot reconstruct them afterwards. Two tokens with nearly identical hidden states can land on different experts and produce meaningfully different outputs, a small discontinuity that dense models simply do not have.
ConceptThe idea underneath
Conditional computation rests on a simple economics: quality scales with parameter count, serving cost scales with the parameters you actually multiply by. If those two numbers can be made different, you win. A mixture-of-experts layer does it by replacing one feed-forward network with E of them and adding a router, a single linear layer producing E scores from the token's hidden state.
Only the top k experts run, typically 2 of 8 or 8 of 256. Their outputs are combined weighted by the router's normalised scores, so the layer computes sum over selected e of gate_e * FFN_e(x). The gate is what makes a discrete choice trainable: gradients flow through the selected experts' weights, so a router that picks a useful expert gets reinforced. Newer designs add always-on shared experts to hold knowledge every token needs, freeing the routed experts to specialise.
The hard part is balance. Routers collapse — a few experts get chosen constantly and the rest starve. Training counters this with an auxiliary load-balancing loss, or, in DeepSeek-V3's approach, with per-expert bias terms adjusted during training so no auxiliary loss is needed. At inference none of that applies. The router is frozen and whatever imbalance it learned is what you serve. When experts are distributed across GPUs, an unbalanced batch leaves some devices idle inside every all-to-all and the step takes as long as the busiest device.
Note carefully that the sparsity is in FLOPs, not memory. Every expert must be loaded and resident. That is why these models are fast and demanding at the same time.
At a glanceSee it
A router picks a few experts per token while every expert stays resident in memory.
The knobsHyperparameters and nuance
- num_experts_per_tokthe top-k, 2 in Mixtral and 8 in DeepSeek-V3. Raising it improves quality and raises per-token compute close to linearly; k of 1 is fastest and noticeably weaker.
- num_local_experts or n_routed_expertstotal expert count E. More experts raise capacity and memory with no change to per-token FLOPs, but make balance harder and leave each individual expert less trained.
- n_shared_expertsalways-active experts in DeepSeek-style designs. They cost unconditional compute and stabilise quality, especially at small k, by holding what every token needs.
- scoring_func and norm_topk_probsoftmax over all experts versus per-expert sigmoid, and whether selected gates are renormalised to sum to 1. Model-specific; getting either wrong in a re-implementation yields a subtly worse model and no error.
- expert parallel sizevLLM's
--enable-expert-parallelshards experts across GPUs rather than replicating them, trading an all-to-all per layer for memory. Throughput then becomes sensitive to routing balance. - capacity factorin capacity-limited implementations, the per-expert token cap as a multiple of the average. Below about 1.0 tokens get dropped; well above it, padding wastes compute on nothing.
EffectHow this stage moves the answer
Two answer-level effects, one good and one uncomfortable. The good one: at fixed serving cost an MoE beats a dense model, so the answers you get inside a given latency budget are those of a far larger model — quality tracks total parameters while speed tracks active ones, and that gap is the whole reason the frontier moved this way. The uncomfortable one is discontinuity. Routing is a hard top-k over continuous scores, so a small input change can flip which experts run, and a flipped expert is not a small perturbation, it is a different function. This shows up as brittleness dense models do not exhibit: rephrasing a prompt trivially, adding a trailing newline, or editing an unrelated earlier turn can produce a qualitatively different answer rather than a slightly different one. Under capacity-limited serving it worsens, because whether your token is dropped depends on who else is in the batch.
EvalsWhat it does to your measurements
The first rule is stating which parameter count you are comparing. An MoE at equal total parameters loses to a dense model; at equal active parameters it wins; at equal serving cost it wins by a lot. Benchmarks reporting a single parameter count are uninterpretable for these models, so always give both. The measurement that breaks is latency variance. Expert load depends on the token mix in the batch and the all-to-all step waits for the slowest device, so p99 latency for an MoE deployment moves with workload composition in a way a dense model's does not, and a benchmark on homogeneous prompts will understate it badly. If the serving path enforces expert capacity, add a determinism caveat too: token dropping is batch-dependent, so the same prompt can produce different completions at temperature 0 purely from what it was batched with.
Failure modesWhen it goes wrong
- Tail latency spikes while GPU utilisation stays moderaterouting imbalance leaving some expert-parallel devices idle inside each all-to-all while others queue.
- Memory allocation fails despite a small active-parameter countevery expert is resident; active parameters describe compute and never footprint.
- Output quality degrades under load but not in isolationcapacity-limited routing dropping tokens once the batch fills the per-expert cap.
- A trivial prompt edit produces a completely different answera router score crossed a top-k boundary and a different expert ran.
- A re-implementation is fluent but measurably weaker than the referencewrong gate normalisation or softmax-versus-sigmoid scoring. The model still works, it is just no longer the model that was trained.
PapersWhere this comes from
- Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts LayerShazeer et al., 2017. arXiv:1701.06538. Introduced the sparsely-gated MoE layer with top-k routing and a load-balancing loss, the direct ancestor of every MoE now in production.
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingLepikhin et al., 2020. arXiv:2006.16668. Established expert parallelism with all-to-all dispatch and the capacity-factor mechanism that decides which tokens get dropped under load.
- Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient SparsityFedus et al., 2021. arXiv:2101.03961. Showed top-1 routing is sufficient and can be made stable, and quantified the quality-per-FLOP advantage over dense models.
- Mixtral of ExpertsJiang et al., 2024. arXiv:2401.04088. The open checkpoint that made the serving arithmetic concrete: 46.7B parameters resident, 12.9B active per token, top-2 of 8 experts per layer.
What changedWhat changed here
Updated this page Quantizing the KV cache can flip which experts a MoE model routes tokens to, because routing is discontinuous.
On sub/pipe-mixture-experts-routing, note that MoE routing is discontinuous and that quantizing the KV cache can flip which experts fire.
Three kinds of claim, strongest first. Signal runs every morning.