The number nvidia-smi calls utilization measures whether any kernel was running, not whether the chip was busy. A decode step that saturates HBM can read as thirty percent. The right dashboard has four other numbers on it.
nvidia-smi shows 30 percent utilization on our serving fleet. Is that a problem, and what would you look at instead?
The number nvidia-smi calls utilization measures whether any kernel was running, not whether the chip was busy. A decode step that saturates HBM can read as thirty percent. The right dashboard has four other numbers on it.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on explaining what the utilization counter actually measures, on the bandwidth argument for why decode looks idle, and on naming the DCGM fields and engine metrics that say whether the fleet is actually saturated.
No comments yet — be the first to share your approach.
