← 🧮 Napkin Math & Capacity
Advanced
Cost per Million Tokens
The unit every serving decision cashes out in. It is one formula: the fleet's dollars per second divided by the tokens per second it sustains, scaled to a million, with utilization in the denominator because idle replicas still cost money. This page derives it from a GPU price and a throughput estimate, works it at three batch sizes to show why batching is the main lever, separates prefill from decode pricing, and shows how the same fleet's cost per token moves by 5x between a quiet hour and a busy one.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
LESSONS THAT TEACH THIS
PRACTICE THIS IN REAL QUESTIONS
Napkin Math, Cost & CapacityWhat does it cost per million output tokens to serve a 70B model on eight H100s?→Napkin Math, Cost & CapacityWhat does one training token cost?→Napkin Math, Cost & CapacityRank the H100, MI300X and Trainium2 by cost per token for decode→Napkin Math, Cost & CapacitySize an inference fleet for a 70B model serving 1,000 concurrent users→Napkin Math, Cost & CapacityTraffic peaks at three times the daily average. Capacity-plan the serving fleet.→Napkin Math, Cost & CapacityHow much does moving from bf16 to fp8 save in serving cost?→
