AI evaluation / scientific reasoning
Physics Experts Reevaluate Six AI Benchmarks; GPT‑5.6‑Sol Gains More Than 50 Points After Problems Are Corrected
A human audit covering six physics benchmarks found that most answers initially marked wrong actually involved incorrect reference solutions, grader flaws, or incomplete problem statements. After problematic items were corrected or removed, GPT‑5.6‑Sol achieved 94.4% pass@4 on the retained CritPt problems—but that is not equivalent to its score on the original full leaderboard.

A new preprint brought together faculty members and graduate students from multiple physics subfields to reexamine six benchmarks: HLE-Physics, CMT-Benchmark, CritPt, UGPhysics, PRISM-Physics, and PHYBench. The audit covered the questions, official reference solutions, model responses, and automated grading results. It classified failures as genuine reasoning errors, incorrect reference answers, grader misjudgments, or ambiguous and underspecified questions. The authors report that most of the initial errors they examined fell into the latter three categories, rather than reflecting a genuine inability of the models to solve physics problems.
The largest impact was seen for GPT‑5.6‑Sol: its HLE-Physics mean@4 rose from 47.3% to 78.7%, while its CMT-Benchmark score increased from 61.0% to 87.2%. On CritPt, it originally recorded a mean@5 of 32.3% across 70 research-level challenges. After experts corrected answers and excluded invalid questions, 54 problems remained, on which the model achieved 87.5% mean@4 and 94.4% pass@4. This shows that exact-answer matching and specialized graders are not necessarily reliable: symbolic equivalence, unit conventions, approximation assumptions, and missing information in a question can all cause valid solutions to be marked incorrect.
For model developers, the conclusion is not that “physics reasoning has been solved,” but that the measurement ceiling of many closed-ended scientific benchmarks is already constrained by data quality. Evaluation pipelines should preserve complete reasoning traces and grading records, allow review by domain experts, and report results for the original full dataset separately from those for the corrected subset. Otherwise, a single round of data cleaning can produce what looks like a generational leap in model capability.
The 94.4% result cannot directly replace the CritPt leaderboard score published by Artificial Analysis because the question set, number of retries, and metric all changed. Moreover, the study primarily focused on text-based problems with verifiable final answers. Genuine research capability also includes modeling, formulating hypotheses, designing experiments, and exercising judgment on unresolved questions. The next generation of benchmarks will still need prevalidated problems, versioned answers, and auditable graders.