Twenty models of different sizes is a packing problem with a cold-start penalty, and the two pull in opposite directions. What stays resident, what loads on demand, the arithmetic that decides which is which, and the interface that stops every team from asking for a dedicated replica.
Design a platform that serves twenty open-weights models of varying size to internal teams.
Twenty models of different sizes is a packing problem with a cold-start penalty, and the two pull in opposite directions. What stays resident, what loads on demand, the arithmetic that decides which is which, and the interface that stops every team from asking for a dedicated replica.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on the resident-versus-on-demand decision from request rate and load time, on cold start as the dominant cost, and on the interface that shapes demand.
No comments yet — be the first to share your approach.
