Latency has three numbers and reporting one of them hides the problem
Time to first token, time per output token and end-to-end latency measure three different things, and a change that improves one routinely damages another. Reporting a single average hides both the phase that broke and the users it broke for.
13 MIN · PLUS
a free account unlocks the core curriculum tier · no card
