A driver upgrade touches every layer at once: kernel module, CUDA runtime, container toolkit, NCCL, fabric driver, and every job's image. The compatibility matrix that decides whether a job can run on the new node, the canary that proves it, the wave arithmetic for 2,000 nodes, and the rehearsed rollback.
Upgrade the GPU driver across 2,000 live nodes without breaking running jobs. Walk me through the plan and what can go wrong.
A driver upgrade touches every layer at once: kernel module, CUDA runtime, container toolkit, NCCL, fabric driver, and every job's image. The compatibility matrix that decides whether a job can run on the new node, the canary that proves it, the wave arithmetic for 2,000 nodes, and the rehearsed rollback.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the compatibility matrix (driver, CUDA, container toolkit, NCCL, fabric driver, images), on canary with a real job and measured collectives, on the drain-based rollout arithmetic, and on rehearsed rollback and the kernel-pinning trap.
No comments yet — be the first to share your approach.
