AI Infra Interviews logo
LEARNING PATH

Inference performance

Make a model serve more users per GPU without breaking the latency you promised.

WHO IS ON IT

You want to own a serving fleet: throughput, time to first token, cost per token, and the scheduler decisions behind all three.

WHAT THE LOOP WEIGHTS

Serving mechanics and capacity arithmetic, a design round on batching and the KV cache, and a debugging round where p99 moved and p50 did not.

The course sequence

In this order. Each assumes the one before it.

The question tracks to drill

91 questions across 3 tracks, in the order this loop weights them. Design an inference platform, a batching system, a training cluster, with TTFT and TPOT numbers on the board.

Who hires for this

Grouped by the kind of employer, because archetype predicts the loop better than the brand does.

Keep these open while you work