Model Families & Architectures · 6 min
Mixture-of-Experts: Total vs Active Parameters
Sparse routing sends each token to a few experts of many, making an MoE compute-cheap per token yet still full-price in memory to hold every expert.
A dense transformer runs every parameter for every token. A 70-billion-weight model puts all 70 billion to work predicting each next token. Capacity is welded to per-token cost, and that welding is exactly what Mixture-of-Experts (MoE) cuts. The trick: build a layer out of many parallel sub-networks ("experts") and, for each token, run only a couple of them. You get the representational capacity of a huge model at the compute bill of a small one. The one thing you don't escape is memory. Every expert sits in VRAM whether or not this token touches it.
The gap between the parameters you store and the parameters you compute with is the whole lesson. Hold onto it.
Where the experts live
MoE doesn't replace the transformer. It swaps out one block. A transformer layer is, roughly, attention followed by a feed-forward network (FFN). Attention mixes information across tokens; the FFN is a per-token nonlinear transform, and in most models it holds the majority of the weights. MoE replaces that single FFN with N copies of it, the experts, plus a small gating network (the "router") that decides which experts each token visits.
┌─────────── one MoE layer ───────────┐
token x ──►│ router (a linear layer + softmax) │
│ picks k of N experts │
│ │
│ E1 E2 E3 E4 ... EN │ ← each Ei is a full FFN
│ ▲ ▲ │
│ │ token routed here only │
│ g2·E2(x) + g4·E4(x) ──────────────┼──► output
└──────────────────────────────────────┘
Attention, embeddings, layernorms: SHARED, run every token.The router is tiny: usually one matrix that maps the token's hidden vector to N logits, one score per expert. Softmax those, pick the top k, and combine the chosen experts' outputs weighted by their gate values. Everything outside the FFN (attention, embeddings, norms) stays dense and runs for every token. That detail drives the parameter arithmetic below.
Two routing regimes worth knowing cold
Switch Transformer (Fedus, Zoph & Shazeer, 2021) took the aggressive position: route each token to exactly one expert. Top-1. They called the block a "switch layer" and argued that the earlier orthodoxy, route to at least two so gradients can compare experts, was unnecessary. Top-1 halves the routing compute and the cross-device traffic, and it trained stably all the way up to Switch-C at 1.571 trillion parameters, with each token still touching a single expert's worth of FFN weights. The same team reported up to 7x pre-training speedups at matched compute against a dense baseline.
Mixtral 8x7B (Jiang et al., 2024) is the one most practitioners actually run. Eight experts per layer, router picks the top 2. The output is the softmax-weighted sum of those two:
logits = router(x) # 8 numbers i, j = top2(logits) # e.g. experts 3 and 6 w = softmax(logits[i], logits[j]) # renormalize over the 2 winners y = w[i] * E3(x) + w[j] * E6(x)
The softmax runs over only the two winners, not all eight. The other six contribute nothing and cost nothing.
The parameter arithmetic, done honestly
Here the name misleads people. "8x7B" is not 56B. Mixtral holds about 47B total parameters but activates about 13B per token. Both numbers need an explanation.
Why 47B and not 56B: the "7B" experts share everything that isn't the FFN. Attention, embeddings, and norms exist once and get reused across all eight experts. Only the FFN sub-blocks are replicated eightfold. You pay for eight FFN stacks plus one copy of the shared machinery, and that lands around 47B.
Why 13B active and not 14B: same reason from the other side. Per token you run the shared attention and embeddings (dense, always on) plus two of the eight expert FFNs. Add the shared part to two experts' FFNs and you're near 13B of actual matrix multiplies per token.
STORE (VRAM) COMPUTE (per token) Mixtral 8x7B ~47B ████████████ ~13B ███ Dense 47B ~47B ████████████ ~47B ████████████ Dense 13B ~13B ███ ~13B ███
Read that table as the entire value proposition and the entire trap. Mixtral thinks with roughly the FLOPs of a 13B model and roughly the quality of something much larger, but you have to own the VRAM of a 47B model to load it. Sparse activation buys latency and throughput, not memory.
Builder: Size your box for total parameters and your latency budget for active. A 47B MoE in fp16 needs about 94GB just for weights, so it won't fit a single 80GB card before you quantize, yet each token decodes as fast as a 13B. Provisioning for the 13B feel and then eating an OOM on load is the classic first-day MoE mistake.
What actually goes wrong
Load imbalance and expert collapse. Nothing in the loss inherently pushes the router to spread tokens evenly. Left alone it finds a few popular experts and starves the rest, wasting the capacity you paid to store and creating stragglers when experts live on different GPUs. The standard fix is an auxiliary load-balancing loss that nudges routing toward uniform expert usage. It's a small added term. Drop it and training degenerates.
Token dropping. During training each expert gets a fixed buffer, and the capacity factor sets how many tokens it will accept per batch. When a popular expert overflows, the extra tokens are dropped: they skip the FFN and ride the residual connection through unprocessed. Set capacity too low and you silently mangle tokens; too high and you burn memory on padding. It's a real knob, not a footnote.
Lumpier batching. Because tokens in one sequence scatter to different experts, an MoE's memory-access pattern is less regular than a dense model's. Throughput depends on how well your serving stack groups tokens per expert, and naive implementations leave a lot on the floor.
Defender: MoE widens the attack surface in a quiet way. The router is an input-conditioned control path, so crafted inputs can bias which experts fire, and shared expert buffers mean one request's routing can, in poorly isolated batches, affect another's throughput. That is a routing-contention side channel. Fingerprinting a served model through expert-activation timing is an active research thread. Treat routing behavior as part of the threat model, not an implementation detail.
Researcher: The clean scaling knob is N (expert count) against k (experts per token). Switch pushed N up with k=1; Mixtral kept N modest with k=2 for a smoother quality-versus-imbalance tradeoff. The open question is how much of MoE's win is genuine specialization versus a bigger, sparsely-regularized parameter bank. Mixtral's own routing analysis found that expert assignment tracks surface token identity and position more than any clean semantic domain.
The one line to keep
MoE splits two numbers that dense models fuse: total parameters (what you store, what sets quality and VRAM) and active parameters (what you compute, what sets latency and FLOPs). Every MoE tradeoff — the OOM on load, the balancing loss, the dropped tokens — falls out of that split. When you read "Mixtral 8x7B," translate on the spot to "47B stored, 13B thinking."
This is also why quantization (an earlier module) bites MoE harder than dense models in practice. You're squeezing 47B of weights into the box you were tempted to size for 13B.
Sources
- Fedus, Zoph, Shazeer (2021), Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, arXiv:2101.03961.
- Jiang et al. (2024), Mixtral of Experts, arXiv:2401.04088.
- Mistral AI, "Mixtral of experts" announcement, mistral.ai/news/mixtral-of-experts. — https://mistral.ai/news/mixtral-of-experts/
- Shazeer et al. (2017), Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer, arXiv:1701.06538 (origin of top-k gating and the load-balancing loss).