The AI infra concept map
150 concepts across 13 tracks, drawn with the 688 links between them. 354 of those links (51%) cross tracks, which is the part worth looking at: the concepts that decide AI infra interviews are rarely the ones that sit neatly inside one topic. Hover a concept to see only what it touches. Click to read it.
Where the map is dense
The most-connected concepts are the ones the rest of the library keeps reaching for, which makes them the highest-leverage things to be solid on: NVLink, NVSwitch and PCIe (18), Model Memory Footprint (17), Capacity Planning and Utilization (17), GPU Memory Hierarchy (16), and Collective Communication Primitives (16). If you are deciding where to spend a week, start with a hub rather than a leaf.
Every concept, by track
GPU & Accelerator Architecture12
- NVLink, NVSwitch and PCIe18
- GPU Memory Hierarchy16
- Roofline Model15
- Memory-Bound vs Compute-Bound Kernels15
- Tensor Cores and Matrix Units14
- GPU Generations: A100 to Blackwell13
- Numerics: FP32, BF16, FP8 and FP49
- GPU Execution Model8
- TPU Architecture and Systolic Arrays7
- Trainium and Inferentia6
- AMD Instinct and ROCm6
- Cerebras, Groq and Dataflow Accelerators6
Hardware & Cluster Build-Out11
- NVLink Domains and the NVL72 Rack10
- The Bill of Materials for a Training Cluster10
- Scale-Out Fabric Choice: InfiniBand XDR vs Spectrum-X8
- Accelerator Selection: H100 to B300 and RTX PRO 60007
- Rack Power Delivery and Busways7
- Direct-to-Chip Liquid Cooling and CDUs7
- Cluster Bring-Up: Firmware, Drivers and the Stack7
- Burn-In and Acceptance Testing7
- SXM, PCIe and Rack-Scale Form Factors6
- Cables, Transceivers and the Optics Power Budget6
- Colocation, Power Contracts and Site Selection6
Distributed Training13
- Collective Communication Primitives16
- Tensor Parallelism13
- Expert Parallelism for MoE11
- ZeRO and FSDP10
- Elastic and Fault-Tolerant Training10
- MFU and HFU9
- Checkpointing and Resumption at Scale9
- Ring vs Tree All-Reduce8
- Pipeline Parallelism and the Bubble7
- Activation Checkpointing7
- Data Parallelism and DDP6
- Context and Sequence Parallelism6
- RL Post-Training Infrastructure6
Inference & Serving14
- Inference Autoscaling and Cold Starts16
- Serving Engines: vLLM, SGLang and TensorRT-LLM15
- Continuous Batching13
- PagedAttention11
- Prefix Caching and KV Reuse11
- Quantization for Inference11
- Latency Metrics: TTFT, TPOT and Goodput11
- Prefill vs Decode10
- The KV Cache10
- Chunked Prefill10
- Disaggregated Prefill and Decode9
- Attention Variants: MHA, GQA, MQA and MLA9
- Speculative Decoding7
- Multi-LoRA Serving6
Open Weights & Serving Engines11
- vLLM Server Arguments That Matter11
- Open-Weights Models of 202610
- Reading config.json to Size a Model You Have Never Run9
- Weight Formats: FP8 Blocks, MXFP4 and AWQ8
- Multi-Node Serving Topologies8
- Serving Benchmarks That Do Not Lie8
- Capacity Planning for Open-Weights Fleets8
- Multi-Head Latent Attention and Sparse Indexers7
- SGLang Server Arguments That Matter7
- Expert Parallel and All-to-All Backends7
- Model Onboarding: From Hugging Face to Production7
Napkin Math & Capacity11
- Model Memory Footprint17
- Capacity Planning and Utilization17
- Bandwidth-Bound Decode Throughput15
- Cost per Million Tokens13
- TCO: Buy vs Rent13
- KV Cache Sizing10
- Power and Datacenter Constraints9
- GPU-Hours and Time to Train8
- Communication Volume Estimates8
- Arithmetic Intensity by Operation7
- Training FLOPs: 6ND6
Networking & Storage11
- NCCL and Collective Algorithms11
- Parallel Filesystems vs Object Storage11
- Checkpoint I/O11
- Rail-Optimized and Fat-Tree Fabrics10
- Data Loading Pipelines for Training10
- Debugging a Slow All-Reduce10
- RDMA, InfiniBand and RoCEv29
- GPUDirect RDMA and GPUDirect Storage9
- Topology-Aware Communication9
- Congestion Control for AI Fabrics7
- Dataset Lifecycle: Ingest, Shard and Retain6
Scheduling & Orchestration11
- Multi-Tenancy, Quotas and Fair Share12
- Spot, Preemption and Capacity Strategies12
- Topology-Aware Scheduling10
- Containers, Images and GPU Cold Starts10
- Gang Scheduling with Kueue and Volcano9
- Kubernetes GPU Scheduling8
- MIG, MPS and Time-Slicing7
- Slurm for AI Clusters7
- Slurm vs Kubernetes7
- Ray on Kubernetes6
- Node Lifecycle: Drain, Upgrade and Return6
Fleet Reliability & Observability11
- Incident Response for GPU Fleets15
- Node Health Checks and Burn-In13
- DCGM and GPU Telemetry12
- Stragglers and Hangs10
- Training Uptime and Interruption Statistics10
- SLOs for AI Systems10
- GPU Failure Modes and XID Errors9
- NVLink and Fabric Faults9
- Thermal, Power and Cooling Events9
- ECC, Row Remapping and Memory Errors7
- Alert Design and On-Call Load6
AI Systems Design12
- Capacity and Backpressure15
- Inference Platform Architecture14
- GPU Job Scheduler Design14
- Request Routing and Load Balancing for LLMs11
- Training Cluster Design at 10k GPUs11
- Designing for Latency SLOs11
- The AI Infra Design Round Playbook8
- Evaluation and Data Pipeline Infrastructure7
- Serverless GPU Platforms6
- Multi-Tenant Fine-Tuning Service6
- Control Plane and API Design for GPU Platforms6
- Multi-Region Serving and Failover6
Coding for Infra11
- The Practical Coding Screen Playbook11
- Rate-Limiting Algorithms9
- Batching Queues and Backpressure9
- The GPU Credit Scheduler Pattern8
- Concurrency in Python, Go and C++8
- Retry, Backoff and Idempotency7
- Interval Merging and Utilization Logs7
- Producer-Consumer Pipelines7
- Parsing Kernel Traces and Logs7
- Cache-Friendly Data Structures7
- Consistent Hashing and Sharding6
Ownership & Judgment11
- Leveling Signals: Senior vs Staff11
- The Reliability Pushback Story10
- On-Call Narratives That Land10
- Working with Researchers10
- Migrations and Deprecations8
- Deciding Under Incomplete Information8
- Escalation That Works8
- Safety and Mission Rounds at the Labs6
- Your View on Where AI Infrastructure Is Going6
- Mentoring and Growing Engineers6
- Talking About Cost and Capacity with Leadership6
The map is for orientation. If you would rather be told what to do in order, the start-here path sequences the same material by background and stage, and the courses walk it front to back.
