← 🗂️ Scheduling & Orchestration
Advanced
Spot, Preemption and Capacity Strategies
Spot and preemptible GPUs cost a fraction of on-demand and can be taken back with a couple of minutes' notice, so using them well is an expected-value calculation: the discount against the work lost per preemption, which is set by checkpoint cadence and restart time. The same arithmetic governs internal preemption in a shared cluster. This page works the break-even, the checkpoint interval that makes spot pay, and the fleet mix (reserved baseline, on-demand headroom, spot for tolerant work) that a capacity strategy is built from.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
LESSONS THAT TEACH THIS
PRACTICE THIS IN REAL QUESTIONS
Kubernetes, Slurm & GPU SchedulingWe rent GPUs. When should we buy committed capacity instead of paying on demand, and what do we do with the rest of the demand?→Kubernetes, Slurm & GPU SchedulingYour fine-tunes run on spot GPUs preempted about once every four hours. How often should they checkpoint, and when does spot stop paying?→Kubernetes, Slurm & GPU SchedulingOne fleet: training that wants every idle GPU, and inference with a p99 SLO. Separate pools, or one pool with preemption? Show the numbers.→GPU Fleet Reliability & ObservabilityAt 16,384 GPUs something fails every three hours. How often should you checkpoint, and what goodput does that leave?→Behavioral & OwnershipTell me about a cost reduction you led. How did you prove reliability did not suffer?→Kubernetes, Slurm & GPU SchedulingEight research teams share 1,024 GPUs. Design the quota and fairness policy, and tell me how they will game it.→
