Three things change at once: dense compute more than doubles with an fp4 tensor core, HBM3e reaches 8 TB/s, and NVLink stops at 72 GPUs instead of 8. Work through what each does to tensor-parallel degree, per-step weight reads and expert placement for a 671B-parameter mixture of experts.
What actually changes with Blackwell and the NVL72 rack, and what does it do to how you would serve a large MoE model?
Three things change at once: dense compute more than doubles with an fp4 tensor core, HBM3e reaches 8 TB/s, and NVLink stops at 72 GPUs instead of 8. Work through what each does to tensor-parallel degree, per-step weight reads and expert placement for a 671B-parameter mixture of experts.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on separating the three changes, quantifying each against H100 with the datasheet, and reasoning about MoE serving as an expert-placement problem that the 72-GPU domain changes qualitatively.
No comments yet — be the first to share your approach.
