Dollars per FLOP from the GPU price and the MFU, times 6N: a 70B training token costs about three quarters of a microdollar, and the whole 15T-token run follows in one multiplication. The chain, the comparison to an inference token, and why the training token is cheaper.
What does one training token cost?
Dollars per FLOP from the GPU price and the MFU, times 6N: a 70B training token costs about three quarters of a microdollar, and the whole 15T-token run follows in one multiplication. The chain, the comparison to an inference token, and why the training token is cheaper.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
The interviewer wants the candidate to build a dollars-per-FLOP figure with MFU inside it, then multiply by 6N, then check the run total against a known GPU-hour figure. The surprise to surface is that a training token costs less than an output token at inference.
No comments yet — be the first to share your approach.
