- Home
- The Thinking Wire
- Same Price Per Token, 2.32x the Bill: the Model Upgrade Nobody Approved
Same Price Per Token, 2.32x the Bill: the Model Upgrade Nobody Approved
A developer pulled two 14-day windows of his own Codex logs, deduplicated them, and compared. First window, GPT-5.5 xhigh: 1,667 sessions, 12.17 billion tokens, 7.30 million tokens per session. Second window, GPT-5.6 Sol xhigh: 1,715 sessions, 28.22 billion tokens, 16.45 million tokens per session. Session count moved 2.9 percent. Tokens per session moved 125 percent. Total usage came out at 2.32x.
The published per-token prices for the two models are identical. Nobody announced a price increase, because there was none.
That measurement comes from Vincent Schmalbach, an independent developer, working from his own logs across three subscriptions. One operator, one workflow, one machine. Treat the specific multiple as his number, not yours. The method is what transfers, and the method is embarrassingly simple: hold the window length constant, hold the workflow constant, and divide tokens by sessions.
The denominator moved while the price sheet stayed frozen
Most AI cost governance watches the wrong side of the multiplication. Price per million tokens is published, versioned, and easy to put in a spreadsheet. Tokens consumed per unit of work is unpublished, unversioned, and changes whenever the provider rotates which model sits behind a tier label.
Schmalbach’s windows caught exactly that rotation. GPT-5.6 Sol became the top xhigh model, the label stayed the same, the price stayed the same, and the reasoning depth behind the label roughly doubled. In his words: “At the same token price, 2.25x the tokens means roughly 2.25x the cost for a similar token mix.”
There is a second effect that a price-sheet review misses entirely. GPT-5.6 Sol adds a cache-write charge that GPT-5.5 did not have. A new billing dimension appeared while every rate held steady. A finance process that diffs the published price of each model against last quarter’s published price will report no change, twice over, while the invoice climbs.
The operational consequence in his case was blunt: three subscriptions that used to survive roughly a week of extreme work now drain in about one day of medium work.
The upgrade nobody signed off on
Change management for AI spend usually assumes a decision. Someone picks a model, someone approves a budget, someone owns the line. Default-model rotation bypasses all three. The tier label is a pointer, and the provider is free to repoint it. Your workflow config still says “xhigh.” Your runtime is now executing a different cost profile.
This is the same failure shape as an agent approving its own overages, covered in agent budget self-approval, with the authority moved one level up. There, the agent decides to keep spending. Here, the provider decides how much reasoning your unchanged request buys. In both cases the human who owns the budget is not in the loop at the moment the cost is determined.
The fix is not a procurement fix. You cannot negotiate your way out of a denominator you do not measure.
The price war makes this less visible, not more
August 2026 was a loud month for AI pricing. Contrary Research documented OpenAI cutting Terra by 20 percent to $2 input and $12 output per million tokens, and Luna by 80 percent to $0.20 input and $1.20 output, roughly three weeks after launch. Anthropic released Claude Opus 5 at half the price of Fable 5 with comparable coding performance. Google priced Gemini 3.6 Flash below Kimi K3 per task.
Sol’s price was left alone. That detail is easy to lose in a rundown of cuts, and it is the one that matters for anyone whose heaviest workloads run on the top tier. The headline says prices are collapsing. The invoice for a reasoning-heavy agentic workflow says otherwise, because the cuts landed on models that workflow does not use, while the model it does use started consuming twice the tokens for the same job.
Contrary also reports that roughly 95 percent of enterprise AI usage still runs on frontier models. If that holds, most organizations captured close to none of those cuts. They stayed on the frontier tier, where the price did not move and the consumption did.
Instrument the ratio the vendor does not publish
Cost per completed task is the right governing unit, and we have argued that elsewhere in the margin inversion. What Schmalbach’s measurement adds is the cheapest usable proxy when task completion is hard to instrument: tokens per session, compared across two equal-length windows.
Four properties make it work.
Equal window length removes seasonality and sprint rhythm from the comparison. Deduplication removes double-counted logs, which otherwise inflate the newer window because retries look like work. Holding the workflow constant means a change in the ratio points at the model, not at the user. And reporting session count alongside the ratio proves the difference came from depth rather than volume, which is precisely what 2.9 percent more sessions and 125 percent more tokens per session demonstrate.
None of that requires vendor cooperation. The logs are already on your side of the boundary.
The trigger to re-run it is any model swap, including the ones you did not initiate. A provider announcement that a new model is now the default for a tier is a budget event. Treat it like a schema migration: measure before, measure after, keep both numbers.
Borrowing Schmalbach’s 2.25x as your planning figure would repeat a mistake we have written about in borrowed thresholds. His workflow is his. Run the same two windows on your own logs and you will get your own multiple, which may be 1.1x or may be 3x depending on how much of your work sits in deep reasoning.
Routing is the lever, measurement is the trigger
Once the ratio is instrumented, runtime model routing becomes something you can actually govern. A tier that doubled its token consumption for a class of tasks that never needed deep reasoning is a routing decision waiting to be made. Without the ratio, routing is a guess dressed as an optimization.
The reverse also holds. If the deeper model finishes tasks in fewer sessions or with fewer retries, 2.25x the tokens might still be the cheaper path per completed task. That is a real possibility and the measurement is what settles it. Schmalbach reported token consumption, not task success rates, so his data cannot answer that question, and neither can anyone’s spreadsheet until they measure both sides.
Do this now
Pull the last 28 days of your AI runtime logs and split them into two 14-day halves. Deduplicate. Divide total tokens by session count for each half. If the ratio moved more than 20 percent, find out which model was serving that tier in each window and check whether any new billing dimension appeared alongside it. Then write down today’s ratio as a baseline, with the model behind it named explicitly, and re-measure the day the provider rotates the default. The price sheet will keep telling you nothing changed. Your own denominator is the only thing that will tell you when it did.
This analysis synthesizes GPT-5.6 Sol xhigh Uses Twice the Tokens of GPT-5.5 xhigh (Vincent Schmalbach, August 2026), a single-operator measurement of his own logs rather than a population study, and The Frontier AI Price Wars Continue (Contrary Research, August 2026), whose pricing figures we present as cited.
Victorino Group helps teams instrument tokens per unit of work and re-baseline on every model swap, so a silent default rotation shows up as a measurement instead of an invoice. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation