AI Infra Interviews logo
LEARNING PATH

Kernel and performance

Reason from a warp to a rack, and make a kernel fast for a reason you can name.

WHO IS ON IT

You want to write and optimise the code the model actually runs on. The deepest path and the narrowest, and it gates hardest on C++ fluency.

WHAT THE LOOP WEIGHTS

A live kernel to optimise or a take-home, profiler output to read, and hardware questions that expect the memory hierarchy from memory.

The course sequence

In this order. Each assumes the one before it.

The question tracks to drill

90 questions across 3 tracks, in the order this loop weights them. Live CUDA or Triton, a kernel take-home, memory coalescing, occupancy, FlashAttention internals, Nsight.

Who hires for this

Grouped by the kind of employer, because archetype predicts the loop better than the brand does.

Keep these open while you work