AI Infra Interviews logo
Behavioral & Ownership / 30
expert★ EssentialNewAnthropicOpenAINVIDIA

Where do you think AI infrastructure is going over the next five years?

Four claims that are defensible from arithmetic available today, each with the counterargument that could sink it, and what each one implies about the work. Dated to 2026, because a thesis with no date is not a prediction.

Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.

Four claims that are defensible from arithmetic available today, each with the counterargument that could sink it, and what each one implies about the work. Dated to 2026, because a thesis with no date is not a prediction.

20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 283 remaining answers · ₹2,000 / $25

The concepts behind this question

Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.

Advanced
🧭 Ownership & Judgment🔒 Premium
Your View on Where AI Infrastructure Is GoingSomewhere in a senior or staff loop an interviewer asks what you think happens next: to GPUs and their challengers, to training at scale, to inference economics, to the tools. It looks like small talk and it is scored. The answer that works is a thesis with a date on it, a reason grounded in numbers you can derive, the counterargument you find strongest, and the thing you would watch to know you were wrong. This page shows how to build such a thesis from the material on this site, gives three worked examples, and lists the answers that sound informed and fail.
Foundational
🖧 Hardware & Cluster Build-Out
NVLink Domains and the NVL72 RackAn NVLink domain is the set of GPUs that can address each other's memory at full fabric speed, and its size is the single most consequential number in a cluster design. Eight on an HGX node, 72 on a GB300 NVL72 rack. Inside the domain a collective moves at terabytes per second over a copper backplane; outside it, the same collective drops to the scale-out fabric at 800 Gb/s per GPU, a gap of roughly twenty times that decides how models are sharded.
Foundational
🖧 Hardware & Cluster Build-Out
SXM, PCIe and Rack-Scale Form FactorsThe same silicon ships in three shapes and the shape decides the deployment. An SXM module is soldered to a baseboard with a full NVLink mesh and needs 700 to 1,400 W of direct power and usually liquid cooling. A PCIe card slots into a standard server, draws through the slot and a cable, and has no NVLink. A rack-scale system like GB300 NVL72 makes the whole rack one NVLink domain and stops being a server at all. Choosing between them fixes your power, cooling, cabling and scheduling story.
Advanced
🧮 Napkin Math & Capacity🔒 Premium
Power and Datacenter ConstraintsThe binding constraint on new GPU capacity in 2026 is not chips or capital but megawatts: an H100 node draws about 10 kW, a GB200 NVL72 rack about 120 kW, and a 100,000-GPU cluster needs on the order of 150 MW with cooling. This page converts GPU counts to power, power to cooling and facility requirements, and both to cost, so a candidate can size a training hall from a power budget and explain why liquid cooling, PUE and the local grid decide where the next cluster goes.
UP NEXT ON YOUR JOURNEY
FEDITOR'S NOTE

Scored on the claims resting on present-day arithmetic rather than speculation, on each carrying its own counterargument, and on the candidate saying what would change their mind.

DISCUSSION · 0

No comments yet — be the first to share your approach.