The cheapest correct answer beats the best answer at scale, so routing is a cost problem with a quality floor. The three routing signals that work, the cascade that trades latency for cost, and the measurement that catches quality loss no latency dashboard shows.
You serve a small, a medium and a large open-weights model. How do you route requests between them?
The cheapest correct answer beats the best answer at scale, so routing is a cost problem with a quality floor. The three routing signals that work, the cascade that trades latency for cost, and the measurement that catches quality loss no latency dashboard shows.
Updated Sep 2026 · Grounded in real AI infrastructure interview loops and written to a senior-engineer editorial bar, with every number worked and every diagram hand-built.
The concepts behind this question
Ranked by how closely each one overlaps this question's topic, so the first card is the thing to read if the answer above moved too fast.
Scored on routing by measured task difficulty rather than by guesswork, on the cascade's latency cost, and on continuous quality measurement per route.
No comments yet — be the first to share your approach.
