Every replacement costs a spare, a maintenance window and a return process; every device kept costs the risk of a job it will fail. Four triggers that decide it without a case-by-case argument, the arithmetic that sets the repeat threshold, and why the device's history beats any diagnostic result.
Write the policy for when a GPU is replaced rather than returned to service. What are the triggers and what do they cost?
Every replacement costs a spare, a maintenance window and a return process; every device kept costs the risk of a job it will fail. Four triggers that decide it without a case-by-case argument, the arithmetic that sets the repeat threshold, and why the device's history beats any diagnostic result.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on written triggers rather than case-by-case judgment, on the cost of both errors, and on device history outweighing diagnostic results for intermittent faults.
No comments yet — be the first to share your approach.
