模型評測
Artificial Analysis Index 4.2 Doubles Private-Test Weighting to 40% and Removes the Saturated GPQA Diamond
The updated index adds benchmarks for agentic knowledge work and long-form PDF reasoning, raising the weight of non-public questions from 20% to 40%. This reduces opportunities to optimize for public benchmark sets, but also makes it harder for external researchers to fully reproduce the rankings.

Artificial Analysis updated its Intelligence Index to version 4.2 on September 4, redefining the capabilities measured by its composite score. The new version adds AA-Briefcase, in which models must handle long-horizon knowledge-work projects designed by industry experts and involving multiple interdependent tasks and thousands of source files. Results are evaluated using both rubric-based scoring and pairwise comparisons. The benchmark assesses not only the final answer, but also verifiable task completion, analytical quality, and presentation quality.
Another new benchmark, GDP.pdf, focuses on single-turn professional document reasoning. Its questions span ten domains, 100 PDFs, and 4,592 pages, requiring models to synthesize body text, tables, charts, footnotes, and exclusion criteria. Answers are scored against 1,275 expert-written atomic criteria, and a result counts toward the headline All-pass Rate only if it satisfies every relevant criterion. By contrast, GPQA Diamond has been removed from the index as leading models’ scores increasingly approached saturation. The answer key and grading prompts for AA-LCR, SciCode’s execution sandbox, and some Elo sampling and anchoring methods were also updated.
The most methodologically significant change is that private, held-out tests now account for 40% of the total score, up from 20%, encompassing AA-Briefcase, AA-Omniscience, and CritPt solutions. This reduces the incentive for model providers to train or tune directly against published questions and makes the leaderboard more representative of unknown workloads. The tradeoff is that external researchers cannot independently inspect question distributions, data contamination, or evaluator error. In other words, resistance to gaming has improved, while transparency and reproducibility have declined.
After the scores were recalculated, Claude Fable 5.1 ranked first overall, followed by GPT-6 Astra. The latter scored 33.2% on GDP.pdf, ahead of GPT-5.6 Sol at 28.2% and Fable 5.1 at 26.2%. These rankings should not be compared directly with earlier scores as a time series because the question sets, weights, and graders have all changed. Procurement or routing systems that use a single Index score as a gate should pin the index version and retain component-level metrics. The next questions are whether v5 will further increase the share of private testing, and how the benchmark operator will address test rotation, leak detection, and cross-version calibration.