How good an answer, for how long you wait

On a phone, quality trades off against time — a model thinks longer to answer better, and a single accuracy number hides that. Give every on-device model a fixed amount of time on an iPhone 17 Pro; each panel ranks them by accuracy (MMLU-Pro + MATH) at that budget. The order changes with time. Hover a model to follow it across all four.

Table — accuracy at each answer-time budget

Each model runs the same 596-item battery once; accuracy at a time budget is that generation truncated to what the model could produce in the time (budget = seconds × device decode tok/s), re-scored with the board's extractors. Greedy decoding is deterministic, so this needs no re-inference. Apple's built-in FM meters no tokens and answers within budget, so its bar is the same at every time — it never thinks longer. MMLU-Pro + MATH only (IFEval can't be truncated). Same data as the board; see methodology.