Triton hands you a block of data and writes the thread-level code for you, which covers most of what a serving or training stack actually needs. The four things it does not give you, the performance you give up in each case with numbers, the reversal as Triton gains Hopper features, and how to decide in the room.
When do you write a kernel in Triton, and when do you have to drop down to CUDA?
Triton hands you a block of data and writes the thread-level code for you, which covers most of what a serving or training stack actually needs. The four things it does not give you, the performance you give up in each case with numbers, the reversal as Triton gains Hopper features, and how to decide in the room.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on knowing that Triton programs are block-level rather than thread-level, on naming specific cases where hand CUDA still wins and roughly by how much, and on treating the boundary as something that moves.
No comments yet — be the first to share your approach.
