Artificial Analysis benchmarks small AI models on iPhone 17 Pro
Artificial Analysis published a report titled 'Benchmarking pocket-scale inference' that compares small AI models on phones, measuring both their intelligence and their end-to-end generation speed. The benchmark defines a 'small' model as one that fits inside 8 GB of memory after quantization, a budget that already includes the memory an 8K-token KV cache needs. The head-to-head intelligence comparisons in the report use a 16K maximum context.
For the inference-speed side of the benchmark, Artificial Analysis partnered with Liquid AI to gather real inference measurements taken on the devices themselves, with results reported for the iPhone 17 Pro. Artificial Analysis states that it has independently validated Liquid AI's inference measurement process.
The report is built around several charts: average intelligence score plotted against end-to-end generation time on the iPhone 17 Pro, model intelligence on its own, token efficiency, and how often models overrun their context budget. A 'Mobile Device Benchmark Set' breaks the intelligence score down by task: BFCL for tool calling, IFBench for instruction following, AA-Omniscience for knowledge accuracy and for the non-hallucination rate, GPQA Diamond for scientific reasoning, and MATH-500 for quantitative reasoning. The report also points to Artificial Analysis' full Intelligence Index pages for models that have been evaluated there, and to a separate methodology page for the complete rules of the benchmark.
The publicly visible text does not itself carry the resulting numeric scores or name which specific small models were tested: the score and time comparisons are rendered as interactive charts rather than as text, and readers are pointed to individual model pages to see them. No publication date or author is given, and although the report refers to 'mobile phones' in the plural, the only device named in the visible material is the iPhone 17 Pro.
Key facts
- Artificial Analysis defines a 'small' model for this benchmark as one that fits inside 8 GB of memory after quantization, including an 8K-token KV cache.
- The intelligence comparisons in the report use a 16K maximum context.
- Real inference measurements were gathered on-device in partnership with Liquid AI, with results reported for the iPhone 17 Pro, and Artificial Analysis says it independently validated Liquid AI's measurement process.
- The Mobile Device Benchmark Set covers six evaluations: BFCL (tool calling), IFBench (instruction following), AA-Omniscience accuracy and non-hallucination rate (knowledge), GPQA Diamond (scientific reasoning), and MATH-500 (quantitative reasoning).
- The visible report text does not include the actual numeric scores or name the individual models tested; those appear only in interactive charts and on separate model pages.
Why it matters
The report combines two things that a single capability score usually leaves apart: how intelligent a small model is on standard evaluations, and how fast it actually runs once quantized and placed on a real phone. By restricting the field to models that fit inside 8 GB after quantization, including the memory an 8K-token KV cache needs, Artificial Analysis compares models under the same constraint a phone actually imposes rather than on paper alone. The inference numbers behind that comparison come from measurements taken on the device itself, gathered in partnership with Liquid AI, and Artificial Analysis says it independently checked that partner's measurement process rather than publishing the numbers on trust.
Who it affects
The comparison is aimed at whoever has to choose a small AI model to run on a phone instead of in the cloud: developers embedding on-device intelligence into a mobile app, and the teams building the small models themselves, since Artificial Analysis says the benchmark set is chosen to represent real-world mobile device usage rather than abstract capability alone. It also touches Liquid AI directly, since its inference measurement process is the one being checked and used as the source of the on-device numbers.
How to use it
The full comparison, including the actual scores, is meant to be read on Artificial Analysis' own site rather than from a static description: the intelligence-versus-speed chart, the token-efficiency and context-overrun figures, and the six-part Mobile Device Benchmark Set breakdown are all presented as interactive charts, with links out to per-model pages for the full Artificial Analysis Intelligence Index scores and to a separate methodology page for the exact rules of the mobile test.
How solid is it
Artificial Analysis frames the inference side of the benchmark as measured rather than estimated: real generation times gathered on physical devices, through a partner process it says it independently validated, rather than a simulated or vendor-supplied figure. What is missing from the visible text is any way to check the resulting numbers directly: no numeric scores, model names, publication date or author appear in the readable material, since the results live in interactive charts and on separate model pages rather than in the text itself.
Risks and caveats
The report refers to 'mobile phones' in the plural, but the only device named in the visible material is a single handset, the iPhone 17 Pro, leaving it unclear from the text how the results would carry over to other phones or to Android hardware. The 8 GB, 8K-context definition of 'small' and the choice of evaluations in the Mobile Device Benchmark Set are also Artificial Analysis' own methodology decisions, laid out on its own methodology page, rather than an industry-wide standard.