- Home
- The Thinking Wire
- 51 Authors Re-Graded Six Benchmarks. The Tests Were Wrong in the Other Direction.
51 Authors Re-Graded Six Benchmarks. The Tests Were Wrong in the Other Direction.
Every benchmark critique we have published argued that the number was too high. In September 2026, a paper with 51 authors, several of them Yale physics faculty, re-graded six physics benchmarks and found the failure running the other way: the tests were marking correct answers wrong.
The paper, submitted 11 September 2026 as arXiv 2609.13009, covers HLE-Physics, CMT-Benchmark, CritPt, UGPhysics, PRISM-Physics and PHYBench. Its finding, in the authors’ own words: “Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models’ physics reasoning.” The abstract gives no percentage. “Most” is the strongest quantifier available, and it is enough to break the instrument.
Three Ways a Test Manufactures a False Negative
The authors name three defect classes: grader errors, incorrect reference solutions, and ambiguous or underspecified questions. None of those is a model failing at physics. Each one is the benchmark failing at its own job, and each produces the same artifact on a leaderboard, a wrong answer that was not wrong.
The scale of the correction is where it gets uncomfortable. For GPT-5.6-Sol, HLE-Physics moves from 47.3% to 78.7% on mean@4, and CMT-Benchmark from 61.0% to 87.2% on mean@4, once the flawed items are removed. Both of those corrected figures are computed on retained evaluation subsets after expert review, which the authors state plainly. The corrected number describes a different population than the original score. A smaller test, graded correctly, produced it. On CritPt, the paper reports 94.4% pass@4 across 54 retained challenges, a different metric again.
Read the direction rather than the magnitude and the conclusion is blunt: “These findings suggest that current benchmarks substantially understate frontier models’ ability to solve well-posed physics problems.” The scoping matters. Well-posed. The audit says nothing about ill-posed problems, which is where we would put most of the work a team actually hands to a model, and the authors do not claim otherwise.
What the Authors Recommend Is Retirement
The forward-looking line in the paper is the part most readers will skip, and it is the one that changes a budget: “Near-saturation on these closed-ended tasks highlights the need for more demanding, expert-validated evaluations.”
That is a research group telling you their own instrument class has reached the end of its useful life for distinguishing frontier models from each other. A patch to the grader would not restore it, because the headroom itself is gone. We argued in benchmark invalidity is not contamination that a broken test is a design defect rather than a leak, and this audit is the strongest evidence for that reading we have seen. The defects were in the grading apparatus, not in the data supply chain.
The Same Month, a Benchmark Went the Other Way
Specific Labs published Real-SWE in September 2026. It runs 8 model-and-harness configurations over 10 tasks drawn from private enterprise codebases, 640 scored rollouts, all at high reasoning. The leading pair, Fable 5.1 with Claude Code, resolves 38.8%. GPT-6 Astra with Codex CLI reaches 33.8%. GPT-5.6 Sol with Codex CLI reaches 16.2%.
The benchmark’s own headline is that “6 of 10 tasks have resolution rates below 15%” and that “No model solves every task.” One task, an analytics stream reducer, scored 0 out of 80 across every model and harness tried.
Hold the caveat while you read those numbers. Real-SWE is vendor-published, by a company with a product in the space, running on licensed private codebases. It is not independently reproducible. What it does offer is task shape: median instruction length of 1,742 characters, and a median reference solution touching 11 files, against 6 for FrontierCode and 6 for DeepSWE. The work is bigger than what the public suites ask for.
The per-task spread is the useful part, because it is not uniform. API keys and environments resolve at 71.9%. A multi-region sweep, 68.8%. Entitlement overage lines, 50.0%. Customer identity migration, 37.5%. Then it collapses: billing schedule migration 14.1%, API token metering 12.5%, S3 datastore measurement 10.9%, linearizable scan 10.9%, tax jurisdiction 3.1%, analytics stream reducer 0.0%. Two of the worst scorers are hard distributed-systems problems, so the spread is a statement about task type, not about regulatory exposure.
Real-SWE also classifies how the agents fail: unverified assumption, missed requirement, integration error, regression, wrong file. Its page states that “Missed requirements are the most common failure” and that “Different models fail in different ways.” Rollout duration does not rescue the picture either. By the benchmark’s own chart counts, 76 of 108 short rollouts failed and 385 of 532 longer ones failed, roughly 70% and 72%. Models fail even in short rollouts.
A Public Score Is Now Uninformative in Both Directions
Put the two together and a public number carries error in both signs at once. The physics audit shows an instrument that manufactures false negatives through its own grading apparatus. Real-SWE shows scores that sit near 38.8% at the top when the work resembles what an enterprise actually maintains. Neither source makes this point alone, and neither can tell you which direction your particular score is wrong in.
We have written that you can win the benchmark and lose the workload, and that the scoreboard is broken at both ends because the score inflates on entry and the cost hides on exit. The September evidence adds the part that removes the last defense. Until now, a leader could treat a public benchmark as a conservative floor, on the reasoning that gaming and contamination push scores up, so the real capability is somewhere below the headline. That defense is gone. A score that can be wrong in either direction, understated by its own grader and inflated by a leak, is not a floor, a ceiling, or a bound.
Do This Now: Change What the Vendor Memo Asks For
Delete the benchmark line from your next model selection memo. It has no remaining decision value, in either direction, and keeping it invites someone to treat it as evidence.
Replace it with the two questions Real-SWE is structured to answer, and which no leaderboard rank can.
Which task class does this pair fail on? Not an aggregate score, a breakdown. In Real-SWE’s data the same suite ranges from 71.9% to 0.0% depending on the task. Your codebase has its own version of that curve, and the tasks your team hands to an agent every week are the only ones that matter. Pick five recent tickets that a model would plausibly own and record resolution per ticket, not an average.
Which failure class does it produce? Missed requirement, unverified assumption, integration error, regression, wrong file. These have different costs and different controls. A missed requirement is caught in review if the requirement is written down. An unverified assumption survives review and fails in production. Knowing which one a given pair produces tells you where to spend review budget, and a benchmark rank tells you nothing about it.
The physics authors asked for “more demanding, expert-validated evaluations.” Inside a company, the expert is your senior engineer and the demanding evaluation is the work already in your backlog. That evaluation is expensive to build and it is the only one that stays honest when the public instruments fail in opposite directions.
This analysis synthesizes How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks (Ali Ansari and 50 other authors, September 2026) and Introducing Real-SWE (Specific Labs (Snagnik Das, Siddhant Paliwal, Janak Sunil), September 2026).
Victorino Group builds task-level evaluations on your own backlog so model decisions rest on your work rather than a public score. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation