Four sections, read in a fixed order, and one number that usually names the bottleneck outright. The report of a kernel at 12 percent of DRAM bandwidth while its memory pipeline reads 82 percent busy, what that gap means, the fix it implies, and the numbers the fixed kernel reports back.
Here is an Nsight Compute report for a slow kernel. Read it, name the bottleneck, and tell me what you would change.
Four sections, read in a fixed order, and one number that usually names the bottleneck outright. The report of a kernel at 12 percent of DRAM bandwidth while its memory pipeline reads 82 percent busy, what that gap means, the fix it implies, and the numbers the fixed kernel reports back.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the reading order (Speed of Light, then Memory Workload, then Occupancy and Warp State), on spotting the memory-pipeline versus DRAM-throughput gap, and on predicting the post-fix numbers rather than just naming a fix.
No comments yet — be the first to share your approach.
