A replica takes a minute to become useful and the traffic changes in seconds. Scale on the wrong signal and you buy GPUs you never use, or you drop requests while they boot. The design is three numbers and two timers.
Design an autoscaler for GPU inference replicas that reacts to load without thrashing.
A replica takes a minute to become useful and the traffic changes in seconds. Scale on the wrong signal and you buy GPUs you never use, or you drop requests while they boot. The design is three numbers and two timers.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on choosing signals that lead the SLO breach, on pricing the cold start against the idle cost with numbers, and on the asymmetric hysteresis that stops the oscillation.
No comments yet — be the first to share your approach.
