One Model, Two Harnesses, 37 Points Apart, and the Better Score Was Cheaper

TV
Thiago Victorino
8 min read
One Model, Two Harnesses, 37 Points Apart, and the Better Score Was Cheaper

ARC Prize ran GPT-6 Astra on ARC-AGI-3 twice. The Standard Harness scored 62.7% and billed $26,098. The Provider Adapter Harness scored 99.9% and billed $18,817. Same model, same benchmark, same tasks. The only thing that changed was what state the scaffold carried between requests, and the number moved 37 points.

Then the cost moved the wrong way. The better run was about 28% cheaper.

What Actually Differed

Greg Kamradt’s write-up defines both harnesses in a single paragraph each, and the difference is narrow enough to fit in a sentence. The Standard Harness “enables a model to carry forward notes it chooses to keep with it throughout the environment.” The Provider Adapter Harness “preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.”

One asks the model to write itself a note. The other keeps the model’s own working state alive across the request boundary. That is the whole delta. No fine-tune, no prompt engineering contest, no different model version.

The efficiency figures on the Provider Adapter runs point at why. They used 49% fewer total tokens and were roughly 3.66x faster by aggregate recorded elapsed time. Our reading of why: a harness where the model must reconstitute its reasoning from whatever notes it chose to keep may be paying for that reconstruction twice, once in tokens and once in accuracy. ARC Prize reports the token, speed and score gaps side by side without claiming a mechanism, so the causal direction here is our inference. Preserve the state and both bills shrink at the same time.

That inversion matters more than the 37 points. Procurement and evaluation teams generally assume a frontier of accuracy against cost, where a better score is bought with more compute. Here the harness change moved both variables in the same direction. The accuracy-cost tradeoff people are used to reasoning about was, in this pair of runs, an artifact of a scaffold choice.

For scale on the other side of the comparison, the same page reports Astra using fewer actions than the human baseline on 96.0% of levels, with 51.7% fewer actions per level on average. The human baseline cost $12.78 per attempted game across roughly 500 participants. Those are the numbers most readers will remember. They are also the numbers that mean nothing without a harness label attached.

The Same Effect Shows Up in Safety Scores

A capability benchmark swinging on scaffold choice is uncomfortable. A safety benchmark doing it is worse, because safety numbers are the ones that end up in disclosure documents and vendor questionnaires.

A researcher writing as richbc, in work done at MATS, built a jailbreak template and evaluated it on ClearHarm’s 179 CBRNE and cyber prompts across 23 models from 7 providers. The template achieves an 84-100% attack success rate on the 9 most vulnerable models. (The post carries an explicit disclaimer that the views are the author’s own, not MATS Research’s, and no real name appears on the page. Treat it as one researcher’s published result, not a lab position.)

The relevant finding for anyone writing an eval policy is what happened when reasoning mode was toggled. Kimi K2.5 dropped from 99.4% ASR to 23.5%. Kimi K3 dropped from 20% to 0%. Gemini 2.5 Flash went the other way, rising from 91.6% to 98.9%.

Three models, one configuration flag, and the direction of the effect is not consistent. A safety score reported without the reasoning-mode setting is not a conservative estimate or an approximate one. It is a number whose sign you cannot infer.

The post’s coverage results carry a second warning. All models returned at least one fully-jailbroken response except Meta Muse Spark 1.1 and more recent Anthropic models (Haiku 4.5, Opus 4.6, Sonnet 5). Claude 3.7 Sonnet, which predates constitutional classifiers, showed 100% ASR. Frontier models, in the author’s words, “are still vulnerable to older jailbreaking techniques if used in combination.” A FAR.AI report quoted in the post’s first footnote makes the methodological version of the point: “Jailbreaks often become more effective when multiple techniques are combined. Robustness evaluation must therefore cover attack compositions rather than only individual jailbreak prompts.”

Attack composition is a harness property. So is reasoning mode. So is state retention between requests. The thing being measured in all three cases is a pair, and only one half of the pair gets printed on the scoreboard.

Why This Breaks Procurement Language

We have argued before that switching evaluation scaffolds moves results, and that the harness is the control surface. The first of those was working with swings around 15%. A 37-point spread on a single benchmark, from one state-handling decision, is a different order of problem.

At 15%, a benchmark citation is imprecise. At 37 points, a benchmark citation without a harness spec is not a weak claim. It has no truth value at all, because the same sentence is simultaneously true and false depending on a configuration detail the reader was never given.

Look at how model capability is actually referenced in the artifacts that bind companies: vendor security questionnaires, model cards summarized into slide decks, AI policy documents that name a threshold score, procurement rules that require a benchmark result above some bar. In the vendor questionnaire responses and AI policy documents we have read, the harness specification is simply not a field. They carry a model name, a benchmark name, and a percentage.

Under those rules, a vendor can satisfy a safety threshold and a capability threshold with entirely legitimate, disclosed runs, while a customer deploying the same model in their own scaffold sees numbers that resemble neither. Nobody lied. The specification was underdetermined.

Do This Now

Open the last document your organization produced that cites a model benchmark score. A vendor questionnaire response, an internal AI policy, a model evaluation report, a board slide. For each score in it, answer two questions in writing.

First: what harness produced this number? Name the state-handling behavior between requests, the reasoning-mode setting, and whether the eval tested single prompts or compositions. If the source did not say, the number is unspecified and your document should say so rather than repeat the figure bare.

Second: is your production scaffold the same one? If your agents carry compacted reasoning state and the cited benchmark ran on a notes-only harness, your deployment is not the configuration that was measured. That is not automatically a problem. It is automatically an unknown, and unknowns belong in the risk register instead of the summary column.

Then change the template. Any internal document that quotes a benchmark should have a required harness field next to the score, empty by default and visibly empty when nobody filled it. A blank field is a governance signal. A confident percentage with no provenance is not.

The cheapest version of this work is renaming a column. The expensive version is discovering, after deployment, that the 99.9% you procured against was the other harness.


This analysis synthesizes GPT-6 Astra on ARC-AGI-3 (ARC Prize, Greg Kamradt, September 2026) and From safety research prompt to cross-model universal jailbreak (richbc, work done at MATS, September 2026).

Victorino Group helps engineering organizations specify and govern the evaluation harnesses behind the scores they act on. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation