Decode throughput is bytes per step over bandwidth, price is dollars per hour, so tokens per dollar for a 70B falls out of a spec sheet in ten lines. In bf16 the three generations are within 10% per dollar; fp8 pulls the H100 ahead, and fp4 on the B200 doubles it again. The chain, the table and the caveats.
Compare A100, H100 and B200 for a 70B serving fleet. Which gives the most tokens per dollar, and where do fp8 and fp4 change the ranking?
Decode throughput is bytes per step over bandwidth, price is dollars per hour, so tokens per dollar for a 70B falls out of a spec sheet in ten lines. In bf16 the three generations are within 10% per dollar; fp8 pulls the H100 ahead, and fp4 on the B200 doubles it again. The chain, the table and the caveats.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on running the bandwidth-bound decode chain for each part at a stated batch and context, converting to dollars per million tokens, and correctly identifying that precision, not the generation, is what moves the ranking.
No comments yet — be the first to share your approach.
