Artificial Analysis releases Intelligence Index v4.2

On September 4, 2026, benchmarking group Artificial Analysis released Intelligence Index v4.2, an interim update that carries over elements it had planned for a coming v5 release. The company says it deliberately held back changes to keep the Index stable through recent major model launches, but decided the pace of frontier releases in the past few weeks made an immediate update necessary. Index v4 launched in January, putting v4.2 about eight months later.
The changelog adds two evaluations and drops one. AA-Briefcase is Artificial Analysis's own in-house evaluation, built on a private held-out test set, that tests agentic knowledge work: multi-week projects built by industry experts, each containing many linked tasks and thousands of input source files. It combines rubric and pairwise grading to score verifiable task success, analytical quality and presentation quality. GDP.pdf, created by Surge AI, evaluates single-turn professional document reasoning across 100 PDFs spanning ten domains; models must synthesize evidence spread across 4,592 pages of text, tables, charts, footnotes and exclusions, and responses are graded against 1,275 expert-authored atomic criteria, with the headline All-pass Rate crediting a task only when every criterion is met. GPQA Diamond, a scientific-reasoning evaluation, is removed; Artificial Analysis says it has become saturated.
Private, held-out test sets now make up 40% of the Index's overall weighting, double the share in v4.1. The held-out pool includes AA-Briefcase, AA-Omniscience, and the solutions for CritPt. Artificial Analysis says this cuts labs' ability to game the evaluations and that the held-out share will grow further in v5. The update also brings grading changes: AA-LCR v1.1 gets a new grading system prompt plus corrections to errors and ambiguities in its answer keys; GDPval-AA v2 and AA-Briefcase have improved sampling and a re-anchored Elo scale for more stable ratings as new models join; and SciCode's grading sandboxes were made more robust so that slow but correct code no longer counts as a failure.
On the results, Anthropic's Claude Fable 5.1 leads the overall Index, followed by OpenAI's GPT-6 Astra, which posts a 4-point gain over its predecessor GPT-5.6 Sol. Meta ranks third, ahead of SpaceXAI, Moonshot/Kimi, Z.AI and Google, in that order. On the Cost per Task Pareto frontier, four labs, Anthropic, OpenAI, Meta and Z.AI, share the frontier. GPT-6 Astra is described as more token-efficient than almost every other model near the intelligence frontier (among models scoring above 25 on the Index), with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite sitting at either end of that efficiency curve. On AA-Briefcase specifically, Anthropic's Claude Fable 5.1 and Opus 5 lead, followed by GPT-6 Astra and Muse Spark 1.3; GPT-6 Astra shows a gain of about 85 Elo points over GPT-5.6 Sol on that evaluation. On GDP.pdf, OpenAI leads: GPT-6 Astra scores a 33.2% All-pass Rate, GPT-5.6 Sol scores 28.2%, and Claude Fable 5.1 scores 26.2%.
Key facts
- Artificial Analysis released Intelligence Index v4.2 on September 4, 2026, an interim update ahead of a planned v5.
- It adds two evaluations, AA-Briefcase (a private agentic knowledge-work benchmark) and Surge AI's GDP.pdf (document reasoning across 4,592 pages), and drops GPQA Diamond as saturated.
- Private, held-out test sets now make up 40% of the Index's weighting, double the share in v4.1.
- Claude Fable 5.1 leads the overall Index; GPT-6 Astra ranks second with a 4-point gain over GPT-5.6 Sol.
- On GDP.pdf's All-pass Rate, GPT-6 Astra leads at 33.2%, ahead of GPT-5.6 Sol at 28.2% and Claude Fable 5.1 at 26.2%.
Why it matters
The Intelligence Index is one of the most widely cited cross-lab model leaderboards, so a change to what it measures reshapes how model capability gets compared in public. Doubling the share of hidden, held-out test sets to 40% directly targets a known weak point of public benchmarks: labs optimizing for tasks they can see in advance. Adding AA-Briefcase and GDP.pdf shifts the Index toward longer, messier, more realistic work (multi-week agentic projects, document reasoning across thousands of pages) rather than short, clean question sets, which is also why GPQA Diamond, now saturated, was retired.
Who it affects
Directly, the labs tracked on the leaderboard: Anthropic, OpenAI, Meta, SpaceXAI, Moonshot/Kimi, Z.AI and Google, whose relative standing shifts with the new task mix and weighting. Indirectly, anyone using the Index to choose a model for agentic knowledge work or heavy document analysis, since AA-Briefcase and GDP.pdf are built specifically around those use cases rather than generic Q&A.
How to use it
The Index, its methodology and per-model breakdowns are published on Artificial Analysis's site for anyone to read; the source names no price or subscription tied to viewing the results. Teams picking a model for agentic project work or long-document reasoning can use the AA-Briefcase and GDP.pdf breakdowns as a closer proxy than the aggregate score alone, since those evaluations map more directly onto that kind of task.
How solid is it
Artificial Analysis describes concrete steps to harden the methodology this round: doubling held-out weighting, adding a grading system prompt and fixing answer-key errors in AA-LCR v1.1, re-anchoring the Elo scale and improving sampling for GDPval-AA v2 and AA-Briefcase, and hardening SciCode's sandboxes so slow-but-correct code stops being marked wrong. That said, it remains Artificial Analysis's own benchmark suite rather than an independently audited one, and the source gives no numeric score for Claude Fable 5.1 on the main leaderboard, only its first-place rank, so the size of its lead over GPT-6 Astra is not stated.
Risks and caveats
The source does not name any individual, spokesperson or executive behind the update, only the company. It also does not give a numeric value for v4.1's held-out test-set share, only that v4.2's 40% is double it; does not state the year Index v4 launched, only that it was eight months before v4.2; gives no reason for GPQA Diamond's removal beyond calling it saturated; and sets no release date for Index v5 beyond saying it is being planned and built. As with any single benchmark suite, the rankings reflect the specific task mix Artificial Analysis chose to weight, not a universal measure of model quality.