Three parallelism dimensions, two link speeds eighteen times apart, and one mapping that makes the job fast. The traffic each dimension generates per step, why the loudest one must stay inside the node, and the arithmetic showing what a wrong rank order costs before anyone notices.
Map a 1,024-GPU job's parallelism onto the hardware. Which dimension goes on NVLink, which on the fabric, and what does a wrong order cost?
Three parallelism dimensions, two link speeds eighteen times apart, and one mapping that makes the job fast. The traffic each dimension generates per step, why the loudest one must stay inside the node, and the arithmetic showing what a wrong rank order costs before anyone notices.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on ranking the dimensions by traffic per step and placing them accordingly, on the NVLink versus fabric bandwidth ratio, and on quantifying the penalty of a wrong mapping rather than calling it suboptimal.
No comments yet — be the first to share your approach.
