- Home
- The Thinking Wire
- The Measured Number Was $12,000. The Rest Was a Baseline Nobody Ran.
The Measured Number Was $12,000. The Rest Was a Baseline Nobody Ran.
OpenAI published a case study claiming $5.9 million in savings from Asana’s migration off Enzyme onto React Testing Library. Gergely Orosz went looking for the measured part of that figure and found about $12,000, spent over roughly two weeks (The Pulse: we need to talk about migrations with AI, August 2026).
Both numbers can be true at the same time. The $12,000 is a cost that was incurred and left receipts. The $5.9 million is a subtraction, and the thing being subtracted from is a project that was never run.
Every savings figure has a measured half and an imagined half
The measured half is what the work cost: tokens, engineer hours, elapsed calendar time. Someone can go check it.
The imagined half is the counterfactual. It is what the same work would have cost under an approach nobody executed. It leaves no artifact, no timesheet, no commit history. A counterfactual is unaudited by construction, and no amount of diligence changes that, because there is nothing on the other side of the diligence to find.
Nothing about this is hidden. The word savings is comparative and presupposes an alternative. The presupposition goes unexamined because a dollar figure reads as a measurement whatever its provenance.
In the Asana case the measured half is $12,000 against a $5.9 million headline. That works out to roughly 0.2 percent of the number a reader takes away. The remaining 99.8 percent is an estimate of a project that does not exist.
Where the $6 million came from
Per Orosz, the baseline was five years and $6 million, built from four engineers at roughly $300K per year. The arithmetic holds up. Four engineers at $300K across five years works out to roughly $6 million, so nothing is being fudged inside the model.
The assumption underneath the model carries all the weight. Orosz’s reading of the estimate is that it likely assumed “a fulltime engineer could do a maximum of X tests migrated per day, where X was between 5 and 10.” That is his inference about the implied throughput, not a rate anyone measured. Move that assumption and the entire headline moves with it. Assume double the throughput and roughly half the claimed savings evaporate, without a single fact about the actual migration changing.
An auditor can verify the multiplication. An auditor cannot verify the assumption, because the assumption has no referent in the world.
The customer walked the baseline back
Asana’s Dan Ubilla later confirmed to The Pragmatic Engineer that the $6 million was “a back-of-the envelope estimation.” It assumed a fully manual migration, with no efficiency gains and no LLM usage factored in, and he called it “probably an overestimation.”
That correction is what makes this case worth keeping. Someone on the customer side, with nothing to gain from the smaller number, said the baseline was too big. Nobody is normally positioned to say that. The figure was published by the vendor and characterised by the customer as back-of-the-envelope, and neither side has an artifact that could settle it. The customer has no reason to revisit it once the case study ships. So the figure travels, uncorrected, into decks and board conversations where it is treated as a measurement.
This is a different failure from the one in the vendor bias in AI productivity metrics. There, the instrument was calibrated by a party with an interest in the reading. Here the instrument is fine and the reference point is fictional. Both produce a number that cannot survive an audit, for unrelated reasons.
What the measured half can actually support
$12,000 and two weeks is a decision-grade number on its own. It tells a buyer the cost of attempting the same class of work, the order of magnitude of the commitment, and how long the team is exposed before finding out whether the approach holds. A CTO can budget against it. If the approach turns out to be wrong, the failure surfaces inside those same two weeks.
The $5.9 million supports no decision at all. It cannot size a budget, because it is not a cost anyone will pay. It cannot compare vendors, because each vendor picks its own baseline. It cannot justify a headcount plan, because the four engineers in that model were never hired. It works as a magnitude signal in a headline and stops there.
The asymmetry is worth sitting with. The small number is the one with evidence behind it. The large number is the one that travels, and the large number is the one that ends up on the slide.
The migrations themselves are real
None of this makes the work unimpressive, and Orosz does not argue that it does. His conclusion runs the other way: AI genuinely makes migrations viable that were previously impractical, and it lowers the burden of supporting legacy systems while a transition is in flight. That is a substantive change in engineering economics, and it does not depend on the $5.9 million being right.
The other two cases he cites have the same shape. Airbnb migrated 3,500 test files in six weeks in 2025, against an estimated 1.5 engineering years of manual effort. Uber migrated 600,000 unit tests across 15 million lines of code in four months with two engineers, modifying 1.25 million lines in the process.
Read those carefully and the split is visible in each one. File counts, test counts, line counts, engineer counts, elapsed months: measured. The 1.5 engineering years Airbnb was comparing against: an estimate of something nobody did. The measured half tells a buyer the work is achievable at that scale, which is the useful signal. The imagined half tells the buyer what someone wants them to believe about the road not taken.
Do this now
The next time a savings figure lands on your desk, whether from a vendor, a consultancy, or your own team, ask one question before anything else:
Which half of this number was measured, and which half is a baseline nobody ran?
Then follow it with the two that make the answer usable. What did the work actually cost, in tokens and hours and calendar time? And who produced the baseline, using what assumed throughput?
If the answer to the first question is a small measured cost and a large imagined one, the figure is not wrong. It is simply two different kinds of claim glued into a single figure, and only one of them has evidence behind it. Cite the measured half. Treat the counterfactual as an opinion held by whoever wrote it, and ask them what their assumed rate was.
Organizations that skip this step end up making real decisions against imagined baselines. That is how a productivity claim becomes a headcount reduction nobody audited, and how a substitution narrative fails its own audit once someone finally checks. The Asana case is the only correction of this kind I have seen published, from the customer, on the record. Most of the time the question has to come from the buyer.
This analysis draws on The Pulse: we need to talk about migrations with AI (The Pragmatic Engineer, August 2026).
Victorino Group helps boards and CTOs separate the measured half of an AI ROI claim from the counterfactual half. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation