Artificial Analysis Intelligence Index v4.2 Released
Original: Artificial Analysis Intelligence Index v4.2
Why This Matters
Benchmark methodology shifts toward private test sets signal a broader industry push to ensure AI evaluations reflect real-world capability.
Artificial Analysis released Intelligence Index v4.2 on September 4, 2026, adding two new evaluations—AA-Briefcase and GDP.pdf—while doubling private test set weighting to 40% to prevent benchmark gaming. Claude Fable 5.1 leads, followed by GPT-6 Astra.
Artificial Analysis has published an interim update to its AI Intelligence Index, version 4.2, ahead of a planned v5 release. The update introduces AA-Briefcase, an in-house agentic evaluation using a private held-out test set that tasks models with multi-week knowledge work projects built by industry experts, graded on task success, analytical quality, and presentation quality. Also added is GDP.pdf, created by Surge AI, which tests long-context document reasoning across 4,592 PDF pages spanning 100 documents and ten domains, graded against 1,275 expert-authored criteria. GPQA Diamond has been removed due to saturation. Private test sets now account for 40% of total Index weighting—double the v4.1 figure—to reduce the ability of labs to game benchmarks. Grading infrastructure has also been upgraded for accuracy and stability. In the current rankings, Anthropic's Claude Fable 5.1 leads overall, followed by OpenAI's GPT-6 Astra, with Meta ranked third. The Cost per Task frontier is shared by Anthropic, OpenAI, Meta, and Z.AI. GPT-6 Astra leads on output token efficiency. Artificial Analysis stated the v5 release has been in development for eight months and that incremental updates will continue.