A router that sends 90% of tokens to 26 of 256 experts makes those ranks do nine times their share while everyone else waits. The arithmetic of the imbalance, what a capacity factor drops on the floor, why the auxiliary loss fights the model, and the bias-only balancing DeepSeek-V3 used instead.
Your MoE router sends 90% of tokens to 10% of the experts. What happens to the step, and how do you fix it without hurting the model?
A router that sends 90% of tokens to 26 of 256 experts makes those ranks do nine times their share while everyone else waits. The arithmetic of the imbalance, what a capacity factor drops on the floor, why the auxiliary loss fights the model, and the bias-only balancing DeepSeek-V3 used instead.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on quantifying the step-time damage, on knowing that capacity limits trade dropped tokens for bounded time, and on explaining why a per-expert bias outside the gradient path balances without distorting the gate.
No comments yet — be the first to share your approach.
