8 TB/s over the bytes per step: 57 tokens per second single-stream in bf16, 113 in fp8, 227 in fp4, and the batch curve on a single card that now holds the whole model. Where the curve bends and what caps it.
How many tokens per second can one B200 decode for a 70B model?
8 TB/s over the bytes per step: 57 tokens per second single-stream in bf16, 113 in fp8, 227 in fp4, and the batch curve on a single card that now holds the whole model. Where the curve bends and what caps it.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
The interviewer wants the candidate to build the batch curve from bytes per step, not quote a single number, and to notice that a B200 holds a 70B in fp8 on one card, which changes the deployment shape.
No comments yet — be the first to share your approach.
