← 🩺 Fleet Reliability & Observability
Advanced
NVLink and Fabric Faults
The links between GPUs are the part of a training node with the most connectors, the highest signalling rates and the least forgiveness: one marginal NVLink cable or one NVSwitch port turns an eight-GPU node into a straggler that slows a thousand-GPU job, and the symptom arrives as an NCCL timeout three layers away from the cause. This page covers what the links are, what their error counters mean, how a fault shows up in NCCL and in step time, how to isolate it to a GPU, a cable or a switch, and the arithmetic of why one degraded link is a whole-job problem.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
LESSONS THAT TEACH THIS
PRACTICE THIS IN REAL QUESTIONS
GPU Fleet Reliability & ObservabilityWhat do NVLink errors look like in telemetry, when is a link degrading rather than broken, and when do you drain the node?→Hardware, Cabling & Cluster Build-OutWhat does nvidia-fabricmanager do, and what exactly breaks when it is not running?→GPU & Accelerator ArchitectureOne node in your cluster shows half the expected NVLink bandwidth in nccl-tests. Walk me through isolating it.→Hardware, Cabling & Cluster Build-OutNew nodes arrive with no software. Walk me through bring-up, and say why the order matters.→GPU & Accelerator ArchitectureMIG versus MPS: what isolation does each give you when sharing a GPU, and which would you pick for a multi-tenant inference node?→GPU Fleet Reliability & ObservabilityWhat is an XID error, which ones mean the hardware is bad, and which ones mean somebody's kernel has a bug?→
