On-Device LLM Leaderboard — intelligence × decode speed, measured under real device limits, with Apple's built-in Foundation Model on the board as sea level.
Every row runs the same 596-item battery on its native runtime, under one fixed on-device budget — cap 4096, greedy, and no-answer scored wrong. The score is what a model actually delivers within those limits; the per-row diagnostics (answered %, median generated tokens) show why a number lands where it does. Cloud APIs on the same battery are a reference ceiling, not ranked.
y = intelligence composite (item-bootstrap mean of IFEval mean-of-4, MMLU-Pro & MATH completed-only accuracy) with 95% CI whiskers. x = decode tok/s, device-measured on iPhone 17 Pro (PipelinedBench, numerics-gated; per-entry gates in proofs). Dashed line = Apple's built-in Foundation Model (no bundle to load). Download size per model is in the table below.
Follows written instructions exactly — e.g. "write 300+ words and never use a comma". Checked programmatically (official checkers); no answer key to memorize.
Expert 10-choice questions across 14 fields (law, medicine, CS, …). Score = % correct of the questions it finished within budget.
Competition math — must produce the exact final answer (symbolically checked). Score = % correct of finished items.
What the model can do, absolute. Different from ③ retention, which is quantized ÷ float of the same model — what shipping it cost.