AI Infra Interviews logo
413 questions · 13 topics · ordered easy to hard

Every AI infra interview question, grouped by topic.

Work through each track in order — they're sequenced the way real loops escalate. The first questions in every topic are free. New here? Start with the AI infra interview questions overview. Sign in to track your mastery as you go.

PRACTICE TESTSThink you know a topic? Take a timed multiple-choice test on it.400 rotating questions · every answer explained · free with an accountStart →

GPU & Accelerator Architecture

30

SMs, warps and the memory hierarchy, tensor cores, the roofline model, FP8 and FP4 numerics, NVLink and HBM generations, and the non-NVIDIA canon: TPU, Trainium, MI300-class, Cerebras and Groq. The hardware physics every other round assumes you know cold.

Hardware, Cabling & Cluster Build-Out

40

Choosing between H100, H200, B200, B300 and RTX PRO 6000; NVLink domains and rack-scale systems; InfiniBand and Ethernet fabrics; the cables, transceivers and optics power nobody budgets; rack power, busways and liquid cooling; bring-up, burn-in and acceptance. The physical layer every cluster rests on, dated to 2026.

CUDA, Triton & Kernel Engineering

30

Coalescing, shared memory and bank conflicts, occupancy, fusion, tiled GEMM, FlashAttention internals, Triton, CUTLASS, Nsight profiling and torch.compile: the live-coding and take-home round at NVIDIA, Fireworks, Together and the labs' performance teams.

Distributed Training & Parallelism

32

DDP, ZeRO and FSDP, tensor, pipeline, context and expert parallelism, collectives and their cost, MFU, activation checkpointing, elastic and fault-tolerant training, checkpoint economics and the RL post-training stack. Owning the training run at cluster scale.

LLM Inference & Serving

30

Prefill versus decode, the KV cache, PagedAttention and continuous batching, chunked prefill, speculative decoding, disaggregated serving, quantization, vLLM, SGLang and TensorRT-LLM, multi-LoRA and routing: hosting open-weight models at a latency SLO and a cost you can defend.

Open-Weights Models & Serving Engines

40

Running the 2026 open-weights frontier: GLM-5.3, Kimi K3 and DeepSeek V4. Reading config.json to size a model you have never run, latent attention and sparse indexers, vLLM and SGLang configuration, expert parallelism and all-to-all backends, weight formats, and the benchmarks that do not lie.

Napkin Math, Cost & Capacity

30

Memory footprints, 6ND, arithmetic intensity and the ridge point, bandwidth-bound decode, communication volume, GPU counts and time to train, cost per million tokens, TCO and buy versus rent. The estimation round almost every AI infra loop includes, with every assumption stated.

Networking, Interconnects & Storage

30

NCCL and the collective algorithms, RDMA, InfiniBand versus RoCE, rail-optimized and fat-tree fabrics, congestion control, GPUDirect, parallel filesystems versus object storage, data loading and checkpoint I/O: the fabric and the disks that decide whether ten thousand GPUs act like one.

Kubernetes, Slurm & GPU Scheduling

30

Device plugins and dynamic resource allocation, MIG, MPS and time-slicing, gang scheduling with Kueue and Volcano, topology-aware placement, multi-tenancy and quotas, Slurm versus Kubernetes, containers and cold starts: the platform round at CoreWeave, Modal, Nebius and every GPU cloud.

GPU Fleet Reliability & Observability

30

DCGM, the XID taxonomy, ECC and row remapping, NVLink faults, stragglers and hangs, thermal and power events, node health checks, SLOs for training and serving, incident response and postmortems at fleet scale. The on-call reality most prep sites skip.

AI Infrastructure System Design

31

Design an inference platform at 10k requests per second, a 10,000-GPU training cluster, a job scheduler with preemption and checkpointing, a serverless GPU runtime with sub-second cold starts, a multi-tenant fine-tuning service, an eval pipeline. The whiteboard round at OpenAI, Anthropic, Baseten and Together.

Coding for Infra

30

Practical builds in Python, Go and C++: a GPU credit scheduler, a rate limiter, a versioned KV store, merging GPU idle intervals, a batching queue, retry with backoff, concurrency under load, parsing a kernel trace. The screens that test whether you can ship infra code in 45 minutes.

Behavioral & Ownership

30

Pushing back on a launch for reliability, the on-call story, the migration nobody wanted, working with researchers, the safety and mission rounds at the labs, and your view on where AI infrastructure is going. The rounds that decide between two technically equal candidates.