A fleet of ten thousand GPUs generates enough hardware events to page someone hourly, and a rota that receives them stops reading them within a month. The three tests an alert must pass to page, the arithmetic of a sustainable rota, and the automation that has to exist first.
You own the on-call rota for a GPU fleet. What is allowed to wake someone at 3am, and what must not?
A fleet of ten thousand GPUs generates enough hardware events to page someone hourly, and a rota that receives them stops reading them within a month. The three tests an alert must pass to page, the arithmetic of a sustainable rota, and the automation that has to exist first.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the three tests (customer impact, human action required, urgency), on the page-volume arithmetic at fleet scale, and on automation as the precondition rather than an improvement.
No comments yet — be the first to share your approach.
