Kernel and performance
Reason from a warp to a rack, and make a kernel fast for a reason you can name.
You want to write and optimise the code the model actually runs on. The deepest path and the narrowest, and it gates hardest on C++ fluency.
A live kernel to optimise or a take-home, profiler output to read, and hardware questions that expect the memory hierarchy from memory.
The course sequence
In this order. Each assumes the one before it.
The question tracks to drill
90 questions across 3 tracks, in the order this loop weights them. Live CUDA or Triton, a kernel take-home, memory coalescing, occupancy, FlashAttention internals, Nsight.
Coalescing, shared memory and bank conflicts, occupancy, fusion, tiled GEMM, FlashAttention internals, Triton, CUTLASS, Nsight profiling and torch.compile: the live-coding and take-home round at NVIDIA, Fireworks, Together and the labs' performance teams.
30 questionsSMs, warps and the memory hierarchy, tensor cores, the roofline model, FP8 and FP4 numerics, NVLink and HBM generations, and the non-NVIDIA canon: TPU, Trainium, MI300-class, Cerebras and Groq. The hardware physics every other round assumes you know cold.
30 questionsMemory footprints, 6ND, arithmetic intensity and the ridge point, bandwidth-bound decode, communication volume, GPU counts and time to train, cost per million tokens, TCO and buy versus rent. The estimation round almost every AI infra loop includes, with every assumption stated.
30 questionsWho hires for this
Grouped by the kind of employer, because archetype predicts the loop better than the brand does.
