A pre-flight suite is a list of tests with a number on each: DCGM health, a GEMM near the fleet median, NCCL at rated bandwidth, NICs at line rate, links active, storage reachable. What goes in the two-minute gate, what waits for the long diagnostic, and why the gate pays for itself.
What do you run on a GPU node before you let a job land on it, how long does it take, and what happens on failure?
A pre-flight suite is a list of tests with a number on each: DCGM health, a GEMM near the fleet median, NCCL at rated bandwidth, NICs at line rate, links active, storage reachable. What goes in the two-minute gate, what waits for the long diagnostic, and why the gate pays for itself.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the suite as a table with expected numbers and fail actions, on distinguishing the per-job gate from burn-in and the acceptance gate, and on quarantine as the only exit for a failing node.
No comments yet — be the first to share your approach.
