The spares number falls out of the failure rate times the repair turnaround, and the cheap items are the ones people forget. The arithmetic for GPUs, the very different arithmetic for transceivers, and the rule that decides whether to hold a part at all.
How many spare GPUs, nodes, cables and transceivers do you hold for a 2,048-GPU fleet?
The spares number falls out of the failure rate times the repair turnaround, and the cheap items are the ones people forget. The arithmetic for GPUs, the very different arithmetic for transceivers, and the rule that decides whether to hold a part at all.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on deriving the pool from failure rate times replacement time, on treating cheap high-count parts differently from expensive ones, and on the correlated-failure buffer.
No comments yet — be the first to share your approach.
