The number comes from one requirement: the gradient reduction has to finish inside the backward pass that produces it. Working that requirement backward gives a bandwidth per GPU, and the answer lands near the port speed the industry ships, which is not a coincidence.
Why do large training clusters provision roughly 400 gigabits per second per GPU rather than more or less?
The number comes from one requirement: the gradient reduction has to finish inside the backward pass that produces it. Working that requirement backward gives a bandwidth per GPU, and the answer lands near the port speed the industry ships, which is not a coincidence.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on deriving the bandwidth requirement from overlap rather than quoting the number, on knowing which parallelism dimension sets it, and on how the answer changes with model size and generation.
No comments yet — be the first to share your approach.
