Understand the concepts before you drill the questions
A structured path through the ideas AI infrastructure and Applied AI loops actually test. Each concept gives you the intuition, a worked example, and the trade-off interviewers probe, then links straight to the real questions where it shows up. Read it like a curriculum, or jump to whatever you are weakest on.
150 concepts across 13 tracks · foundational concepts are free · see them as a map
Begin the curriculum →🧩 GPU & Accelerator Architecture
How an accelerator actually executes: SMs and warps, the memory hierarchy, tensor cores, the roofline, numerics, interconnects, and how TPUs, Trainium and MI300-class parts differ.
🖧 Hardware & Cluster Build-Out
The physical layer: which accelerator for which workload, SXM against PCIe against rack-scale, NVLink domains, scale-out fabrics, cables and optics, rack power, liquid cooling, the bill of materials, bring-up and acceptance.
⚡ Kernels & Compilers
Why kernels are fast or slow: coalescing, shared memory, occupancy, fusion, tiling, FlashAttention, Triton, and the profiler view that explains all of it.
🕸️ Distributed Training
Every parallelism and what it costs: data, ZeRO/FSDP, tensor, pipeline, context, expert, the collectives underneath, MFU, and surviving failures at scale.
🚀 Inference & Serving
From a forward pass to a serving platform: prefill and decode, the KV cache, batching, paging, speculation, disaggregation, quantization and the engines that ship them.
🧮 Open Weights & Serving Engines
Serving the 2026 open-weights models: sizing from config.json, latent attention and sparse indexers, vLLM and SGLang arguments, expert parallelism, weight formats, multi-node topologies, benchmarking and capacity planning.
🧮 Napkin Math & Capacity
The formula sheet as concepts: footprints, FLOPs, intensity, bandwidth bounds, communication volume, cost per token, and how to size a cluster in your head.
🔌 Networking & Storage
The fabric under the collectives: RDMA, InfiniBand and RoCE, topologies, congestion, GPUDirect, and the storage tiers that feed training and checkpoints.
🗂️ Scheduling & Orchestration
How GPUs become a shared platform: Kubernetes device plugins and DRA, MIG and sharing, gang and topology-aware scheduling, Slurm, quotas, containers and cold starts.
🩺 Fleet Reliability & Observability
What breaks in a GPU fleet and how you see it: XID codes, ECC, NVLink, stragglers, thermal, health checks, SLOs, and the incident craft of keeping a training run alive.
📐 AI Systems Design
The building blocks of the design round: routers, schedulers, queues, caches, autoscalers, multi-tenancy and the numbers that make a design credible.
💻 Coding for Infra
The engineering craft infra screens reward: concurrency, backpressure, idempotency, intervals and schedulers, and the patterns behind the practical builds.
🧭 Ownership & Judgment
The half of the loop most engineers under-prepare: reliability pushback, on-call narratives, working with researchers, and the safety and mission conversations at the labs.
