- Home
- The Thinking Wire
- The Same Model Costs Up to 5× More in One Harness Than Another, With Similar Success Rates
The Same Model Costs Up to 5× More in One Harness Than Another, With Similar Success Rates
Twenty-one model-harness pairs, two benchmarks, and one finding a platform team can act on this quarter: the same model reaches a similar success rate at up to 5× the cost depending on which harness runs it. That is the headline of HarnessTax, a September 2026 study from UC Berkeley and Arena (Melissa Z. Pan, Shuo Yang, Negar Arabzadeh, Wei-Lin Chiang, Ion Stoica, Matei Zaharia). It is the first controlled measurement we have seen that treats the harness as a cost variable on its own, separate from the model underneath it.
We have already written about the vendor-controlled multiplier in your token bill, about the harness as the control surface, and about what harness, loop, ops and eval each mean. None of those had a controlled cost comparison across harnesses. This one does, and the numbers change how a harness should be bought.
The setup
Seven models: Claude Fable 5, Opus 4.8, Sonnet 4.6, Haiku 4.5, GPT-5.6 Sol, GPT-5.6 Luna and Kimi K3. Three harnesses: Claude Code, Codex CLI and Pi. Two benchmarks: SWE-bench Lite and Terminal-Bench 2.0. Each pair ran 30 sampled tasks per benchmark, three repetitions each, with a 100-turn cap, the network blocked and the default web tools disabled.
The scope matters as much as the result. Thirty sampled tasks per benchmark is a sampled estimate, and the authors state a limit themselves: both benchmarks are open-source and may have been in training data. Treat the study as a well-instrumented sample, and read every figure below inside that frame.
Success barely moves. Cost does.
In the authors’ words: “The same model can achieve similar success rates at up to 5x costs.” Success stays within ±2% on SWE-bench Lite and within about ±5% on Terminal-Bench 2.0. Small, not zero. Cost is where the spread lives.
The cleanest example is Claude Fable 5 on SWE-bench Lite. In Claude Code it resolved 97.8% of tasks at $1.33 per resolved task. In Codex, 96.7% at $0.89. In Pi, 96.7% at $0.67. Turns per attempt were nearly identical, 15.3 versus 15.4, so turn count does not explain the gap. Roughly double the cost bought about one point of success.
GPT-5.6 Sol on the same benchmark: 77.8% at $1.54 in Claude Code, 74.4% at $0.44 in Pi. About 3.5× the cost for roughly three points. Tom Tunguz, reading the same study in The Harness Margin Opportunity, puts this one pair as a 71% cost reduction: “$1.540 to $0.441 per resolved task, the same model in both cases.” That 71% is one pair, not the aggregate, and the 5× on the HarnessTax page is the widest spread across the 21 pairs rather than a number attached to one model.
Aggregated, the pattern holds. Geometric-mean cost ratios put Claude Code at about 2.0× Pi and about 1.6× Codex on SWE-bench Lite, and about 1.5× Pi on Terminal-Bench 2.0. Tunguz’s chart across the seven models ranges “from 1.1x to 5.1x”. The tax is never zero.
The provider’s harness is not the safe default
The intuitive purchasing rule is that the model vendor’s own harness must be the tuned one. The data does not support it. From the study: “an alternative harness achieves the highest observed success rate in nine of twelve comparisons.” Those twelve comparisons are the six Anthropic and OpenAI models across the two benchmarks. In three quarters of them, a harness the model’s vendor did not build got the best observed result.
Pi is the harness that keeps showing up on the efficient edge. “Pi reaches the Pareto frontier on both benchmarks by providing just four tools: read, write, edit, and bash.” Four tools, and it sits on the frontier for both benchmarks.
Where does the extra cost come from? Part of it arrives before the model does anything. Claude Code’s mean initial context is over 10× Pi’s in tokens. The authors put it directly: “A harness tax can begin with the first model call.” Whatever a harness loads before your task text arrives is billed on turn one, and the study found the two harnesses differ by an order of magnitude at that point.
The statistics make it a governance point
Tunguz added his own analysis on top of the study, and it is the part I would put in front of a CFO. “Across all 42 within-model harness comparisons, a two-sided Fisher exact test finds one result at p < 0.05, where chance alone would produce about two, & none survives a Holm-Bonferroni correction.”
Read as a purchasing rule: no success-rate difference between harnesses survives multiple-comparison correction. Every cost difference is a measured line on the invoice, and the study tests none of them. If you are choosing a harness because it scores higher on a benchmark, the study cannot back that choice at this sample size. If you are choosing one because it costs less per resolved task, the study can.
Tunguz extends this into a margin argument for vendors, with an illustrative pair of companies (38% versus 75% gross margin). By his own footnote that example is illustrative, so I will leave it as one. The measured part is enough.
What the numbers do not say
Three limits on what the study establishes, so a reader takes the right lesson from it.
Pi sat on the Pareto frontier for two open-source benchmarks under a 100-turn cap with no network. Whether it is the right harness for your work depends on whether your tasks need tools beyond read, write, edit and bash.
Claude Code, as configured in the study, spent more per resolved task on these benchmarks, and part of that spend starts in the initial context. A harness that ships more scaffolding may earn it on tasks the study did not sample.
Success within ±2% on SWE-bench Lite and within about ±5% on Terminal-Bench 2.0 is small, not zero. On a task set where every point matters, a few points may be worth the spend. Knowing the price of those points is the whole exercise.
Do this now
Run the Pareto audit on your own tasks before you accept any harness default, including the one your model vendor ships.
Pick one model your team already pays for. Pick two or three harnesses you could realistically run it in. Sample 30 tasks from your own backlog, the kind of tickets your agents actually close, and run each pair three times with the same turn cap and the same network policy. Record two numbers per pair: success rate and cost per resolved task. Plot them. Any point that sits to the right of another at the same or lower success rate is paying a harness tax for nothing you can measure.
Then look at turn one. Pull the initial context each harness sends before your task text arrives. If one harness starts an order of magnitude heavier than another, the tax is being charged before the model reads the ticket, and no prompt engineering downstream recovers it.
Buy the harness as a cost decision. The model decides what is possible. The harness decides what it costs, and this study is the first we have seen put a controlled number on that.
This analysis synthesizes HarnessTax: How Much Does the Harness Matter for Coding Agents? (UC Berkeley / Arena, September 2026) and The Harness Margin Opportunity (Theory Ventures (Tomasz Tunguz), September 2026).
Victorino Group runs harness cost audits on your own task set, so the harness you standardize on is the one your bill can justify. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation