Compute the step time the formula predicts, measure the one you have, and the gap is the budget to explain. The isolation order with one metric per suspect: dataloader wait, SM and tensor-core activity, HBM throughput, time in collectives, and the per-rank spread that says it is one machine.
Your training run is at 60% of the step time you projected. How do you find out whether it is compute, memory, network or I/O?
Compute the step time the formula predicts, measure the one you have, and the gap is the budget to explain. The isolation order with one metric per suspect: dataloader wait, SM and tensor-core activity, HBM throughput, time in collectives, and the per-rank spread that says it is one machine.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on starting from the projected step time rather than from a profiler, on having one metric per suspect, and on checking the per-rank spread before blaming any subsystem.
No comments yet — be the first to share your approach.
