Device plugins count GPUs and know nothing else; Dynamic Resource Allocation lets a pod ask for devices by attribute and share them by claim. What the two models can and cannot express, what DRA changes for MIG, NVLink domains and multi-node gangs, and a migration stance for a fleet that runs both.
Kubernetes device plugins versus Dynamic Resource Allocation: what changes for GPU scheduling, and what would you adopt in 2026?
Device plugins count GPUs and know nothing else; Dynamic Resource Allocation lets a pod ask for devices by attribute and share them by claim. What the two models can and cannot express, what DRA changes for MIG, NVLink domains and multi-node gangs, and a migration stance for a fleet that runs both.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the counting-versus-claims distinction, on concrete things DRA expresses that a plugin cannot (attributes, sharing, structured parameters), and on an adoption stance that keeps the plugin for whole-GPU training pods while DRA earns trust.
No comments yet — be the first to share your approach.
