Allocated is not usable: a cluster full of half-empty nodes has hundreds of idle GPUs that no gang can take. The arithmetic of how mixed job sizes fragment a fleet, the packing policies that prevent it, the defragmentation moves that repair it, and what each costs the jobs already running.
Our cluster is 85% allocated but 64-GPU jobs wait for hours. Explain the fragmentation and what a scheduler should do about it.
Allocated is not usable: a cluster full of half-empty nodes has hundreds of idle GPUs that no gang can take. The arithmetic of how mixed job sizes fragment a fleet, the packing policies that prevent it, the defragmentation moves that repair it, and what each costs the jobs already running.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the free-versus-placeable distinction with numbers, on naming packing (best-fit by node) as prevention and migration or drain-to-consolidate as repair, and on the cost of each defragmentation move to running jobs.
No comments yet — be the first to share your approach.
