← 🧮 Napkin Math & Capacity
Advanced
Capacity Planning and Utilization
Capacity planning for GPUs is deciding how many to have next quarter given that they cost money whether busy or not, that demand arrives in bursts, and that a queue near saturation produces waits that grow without bound. This page works the planning arithmetic for a serving fleet (peak demand, headroom, the p99 penalty of running hot) and a training platform (job mix, queue wait, the value of a shared pool), and gives the queueing intuition that makes 70% look full. The number that decides everything is utilization, and it has a ceiling set by latency, not by hardware.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
LESSONS THAT TEACH THIS
PRACTICE THIS IN REAL QUESTIONS
Napkin Math, Cost & CapacityTraffic peaks at three times the daily average. Capacity-plan the serving fleet.→Napkin Math, Cost & CapacitySize an inference fleet for a 70B model serving 1,000 concurrent users→Napkin Math, Cost & CapacityModel our serving request queue with Little's law. What happens as we approach saturation?→Napkin Math, Cost & CapacityWhat does it cost per million output tokens to serve a 70B model on eight H100s?→Open-Weights Models & Serving EnginesYour product forecasts 12,800 output tokens per second at peak. Size the fleet.→Kubernetes, Slurm & GPU SchedulingWe rent GPUs. When should we buy committed capacity instead of paying on demand, and what do we do with the rest of the demand?→
