You cannot remove one rank from a synchronous job; you replace the node and restart from checkpoint, and the whole craft is making that take three minutes instead of thirty. The detection signals, the drain-restart-quarantine sequence, the spare-pool arithmetic, and where elastic training changes the answer.
A GPU in a running 512-GPU training job is throwing errors. How do you get it out of the job without losing the run?
You cannot remove one rank from a synchronous job; you replace the node and restart from checkpoint, and the whole craft is making that take three minutes instead of thirty. The detection signals, the drain-restart-quarantine sequence, the spare-pool arithmetic, and where elastic training changes the answer.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on saying plainly that a synchronous gang cannot lose a rank (restart is the mechanism), on the automated sequence with a time per step, on the spare-pool derivation, and on knowing when elastic frameworks change it.
No comments yet — be the first to share your approach.
