The Two Dials of AI Cost: Datadog Priced the Trade, OpenRouter Measured the Demand

TV
Thiago Victorino
8 min read
The Two Dials of AI Cost: Datadog Priced the Trade, OpenRouter Measured the Demand

Datadog’s engineering team published the ledger behind more than $1 million per month in AI cost savings in August 2026. The largest single line came from one decision: switching the default model from Opus 4.8 to Sonnet 4.6, a change they describe as “an 8% loss of proficiency in executing Datadog workflows while reducing AI costs by 36.7%,” measured across more than 140 internal evaluations. That one governed trade is worth over $687,000 per month.

The same week’s second data point comes from the demand side. OpenRouter analyzed OpenAI’s July 27 to August 14 discount program and found that daily token usage on Luna grew 13.8x and Terra 5.6x during the promotion, while the un-discounted Sol model grew 1.11x. Cheaper tokens produced dramatically more consumption, the pattern named after Jevons.

Put the two together and you have the complete cost picture most AI budget conversations are missing. Your bill has a supply side and a demand side, and a finance team controls neither by waiting. Falling prices grow usage. Rising capability grows usage. The only lever that moves the bill in a chosen direction is a governed trade-off, priced with evidence and signed off as a decision.

The trade Datadog actually made

We covered Datadog’s governance product roadmap when they announced it. This post is about something different: their internal dogfood ledger, the same genre of evidence we examined when Cloudflare rebuilt its stack on its own primitives. A vendor’s internal numbers carry vendor interest, and Datadog is openly writing about its own products here, Cloud Cost Management and Workflow Automation among them. Read the numbers with that flag raised. They are still the most granular per-mechanism breakdown I have seen published for AI spend.

The Opus-to-Sonnet default switch deserves attention for its structure, not just its size. Datadog did not reach for blunt limits. They ran the candidate model through more than 140 internal evaluations, quantified the capability cost at 8%, quantified the dollar saving at 36.7%, and accepted the trade explicitly. The output of the process is a sentence a CFO and a CTO can both sign: we will give up this much proficiency for this much money.

Compare that with the version of this conversation I keep seeing. The model choice is a default someone set in a config file long ago. Nobody can say what the stronger model buys them, because nobody measured task proficiency on their own workflows. So the conversation collapses into procurement reflexes: negotiate the contract, cap the tokens, hope the next price cut fixes it. None of those move the capability-per-dollar ratio, because none of them measure capability.

The rest of Datadog’s ledger follows the same pattern at smaller scale:

  • Lowering the Claude Code CLI default effort setting from high to medium: over $288,000 per month.
  • Cost alerts that notify individual engineers when their spend spikes: 768 distinct users triggered the alert, and spend dropped more than $150,000 in the first week. Keep the units straight when comparing this figure to the others: it is a 7-day number sitting next to monthly numbers.
  • Headroom, their context optimization layer: cost per user down 27.0% ($156.70 to $114.40), input tokens down 39.3%, output tokens down 35.7%, validated in an A/B test across more than 1,000 engineers for one week.

Every line has the same anatomy. A mechanism, a measurement, a number someone can audit. The evals are the load-bearing element in each case: without them, the Opus-to-Sonnet switch is a gamble on quality, the effort-setting change is a guess, and Headroom is a black box. With them, each is a priced decision.

The dial you do not control

The OpenRouter data explains why none of this happens on its own. During the discount window, Terra and Luna’s combined share of OpenRouter traffic went from 0.7% to 7.8%, and 5.3 of those 7.1 percentage points came out of competitors’ share. Usage did not merely shift; it expanded. A 13.8x jump in daily tokens under a discount is a demand curve speaking clearly.

The caveats matter, and OpenRouter states several themselves. Their post-program observation window is only 6 days against a 19-day program, and they note the story may still change. On July 30, OpenAI cut list prices on top of the 50% program discount, which pushed effective discounts to roughly 90% for Luna and roughly 60% for Terra, so the elasticity was measured against a much steeper price drop than the headline suggests. The retention figures (about 32% of the 100K+ participating customers kept using the models after the program, 18% at or above program pace) count customers, not tokens, so a token-weighted view could look quite different. And OpenRouter’s traffic is a window into one router’s users, not a census of enterprise AI usage.

Even hedged that far, the direction is unambiguous, and it matches what we found in the hidden token bill and in the bill that doubled while the price stayed flat: per-token price and total spend move independently, and often in opposite directions. When the price of a unit of AI work falls, teams find more work worth doing at the new price. That is good news for output. It is fatal for any cost plan built on the assumption that vendor price cuts will flow through to the bottom line.

Two dials, one console

So a CFO looking at an AI line item has exactly two dials that respond to being turned.

The first is the capability-to-cost dial, and it only exists if you have evals. Datadog could accept 8% for 36.7% because they could measure 8%. An organization without task-level evaluations on its own workflows cannot see this dial, let alone turn it. Every model default, every effort setting, every context optimization is a trade of quality against money, and the trade is happening whether or not anyone prices it. Unpriced, it drifts toward whatever the vendor’s defaults happen to be.

The second is the demand dial, and it turns through policy rather than price. Routing rules that send simple tasks to cheap models. Per-engineer visibility that makes spend a felt quantity (Datadog’s alert pilot moved six figures in a week on visibility alone). Effort defaults chosen deliberately. The OpenRouter data shows what happens when this dial is left alone during a price drop: consumption expands to fill the budget, then keeps going.

Waiting is the strategy that turns neither dial. Prices will keep falling, and if the elasticity data generalizes, your bill will not fall with them.

Do this now

Pick the one AI workflow with the largest monthly spend. Build or borrow an evaluation set for it, even twenty representative tasks. Run the current model and the next cheaper model against that set this week. Then write the sentence Datadog wrote: switching costs us X% proficiency and saves Y% of spend. If you cannot fill in X, that is the finding. It means your model spend is an unpriced trade, and the demand curve is the only thing steering it.


This analysis synthesizes How Datadog saves over $1 million each month by optimizing AI usage (Datadog, August 2026) and GPT 5.6 Discounts & Jevons Paradox (OpenRouter, August 2026).

Victorino Group helps engineering organizations build the eval infrastructure that turns AI spend from an unpriced default into a governed trade. Let us talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation