← 🧩 GPU & Accelerator Architecture
Advanced
GPU Generations: A100 to Blackwell
Four NVIDIA generations are in fleets at once, and interviewers ask what each one changed, not what it is called. A100 to H100 added fp8 and tripled compute; H200 kept the die and grew memory; B200 doubled everything and added fp4; B300 stacked more HBM and cut fp64. This page carries the dense numbers for each, what they did to training and serving, and the marketing traps (sparse peaks, 192 versus 180 GB, die counting) that trip candidates. Dated September 2026.
Unlock the full curriculum — ₹2,000 / $25every concept + every answer · 6 months · no auto-renew
RELATED CONCEPTS
PRACTICE THIS IN REAL QUESTIONS
GPU & Accelerator ArchitectureCompare A100, H100 and B200 for a 70B serving fleet. Which gives the most tokens per dollar, and where do fp8 and fp4 change the ranking?→Hardware, Cabling & Cluster Build-OutYour fleet is decode-heavy. Is a B300 worth 1.4 times a B200's power for 1.6 times the memory?→GPU & Accelerator ArchitectureHow much faster is an H200 than an H100, really? Which workloads see the gain and which do not?→GPU & Accelerator ArchitectureWhat actually changes with Blackwell and the NVL72 rack, and what does it do to how you would serve a large MoE model?→Napkin Math, Cost & CapacityHow many tokens per second can one B200 decode for a 70B model?→GPU & Accelerator ArchitectureYou are moving a model to fp4 inference on Blackwell. What breaks first, and how would you measure whether the result is acceptable?→
