A dead node at 4 a.m. costs twenty minutes if a human restarts it and under five if the system does. The four stages of an automatic recovery, the reason a lost node takes a whole pipeline replica with it, why hot spares beat resharding, and the minutes that each stage still costs even when everything works.
Design a training system that survives losing a node without a human in the loop. What does elasticity cost you?
A dead node at 4 a.m. costs twenty minutes if a human restarts it and under five if the system does. The four stages of an automatic recovery, the reason a lost node takes a whole pipeline replica with it, why hot spares beat resharding, and the minutes that each stage still costs even when everything works.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on breaking recovery into detect, reschedule, reshard-or-replace, and resume, on knowing that the parallel layout constrains what can shrink, and on being honest that elasticity is minutes of bubble and a lot of code.
No comments yet — be the first to share your approach.
