A 400B dense model on 10 trillion tokens is 2.4 × 10²⁵ FLOPs, 43 days of pure compute on 16,384 H100s, and about 380 interruptions along the way. The order in which to derive every number, the layout and the data rate, the checkpoint and failure budgets, and the schedule that survives its own arithmetic.
Plan a 10-trillion-token pre-training run end to end: compute, fleet, layout, data, checkpoints and a schedule with a failure budget.
A 400B dense model on 10 trillion tokens is 2.4 × 10²⁵ FLOPs, 43 days of pure compute on 16,384 H100s, and about 380 interruptions along the way. The order in which to derive every number, the layout and the data rate, the checkpoint and failure budgets, and the schedule that survives its own arithmetic.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on deriving each number from the previous one in the right order, on turning the failure rate into days of schedule, and on knowing which numbers are firm (FLOPs, memory) and which are bets (MFU, goodput).
No comments yet — be the first to share your approach.
