AI Infra Interviews logo

Memory decides your parallel layout before you get an opinion

Training state is roughly sixteen bytes per parameter before activations, so most large models do not fit on one device by an enormous margin. What does not fit forces the sharding, and the layout you choose is what remains after that constraint has taken its share.

14 MIN · PLUS

a free account unlocks the core curriculum tier · no card