Generalist AI infrastructure
Serve models and keep the fleet alive, which is what most postings with this title actually mean.
You want the common shape of the job rather than a specialism: enough serving to run it, enough fleet to keep it up, and enough arithmetic to size it.
A broad loop: capacity arithmetic, a serving design round, an operational debugging round, and behavioural questions about owning a shared system.
The course sequence
In this order. Each assumes the one before it.
The question tracks to drill
90 questions across 3 tracks, in the order this loop weights them. GPU credits, dollars per token, utilization, buy versus rent, capacity forecasting.
Memory footprints, 6ND, arithmetic intensity and the ridge point, bandwidth-bound decode, communication volume, GPU counts and time to train, cost per million tokens, TCO and buy versus rent. The estimation round almost every AI infra loop includes, with every assumption stated.
30 questionsPrefill versus decode, the KV cache, PagedAttention and continuous batching, chunked prefill, speculative decoding, disaggregated serving, quantization, vLLM, SGLang and TensorRT-LLM, multi-LoRA and routing: hosting open-weight models at a latency SLO and a cost you can defend.
30 questionsDevice plugins and dynamic resource allocation, MIG, MPS and time-slicing, gang scheduling with Kueue and Volcano, topology-aware placement, multi-tenancy and quotas, Slurm versus Kubernetes, containers and cold starts: the platform round at CoreWeave, Modal, Nebius and every GPU cloud.
30 questionsWho hires for this
Grouped by the kind of employer, because archetype predicts the loop better than the brand does.
