- Home
- The Thinking Wire
- If Agents Are Doing the Spending, They Need a Budget API to Read
If Agents Are Doing the Spending, They Need a Budget API to Read
Software price inflation sat at 16.4% in June 2026, against roughly 2.7% general inflation in the G7. That is the Vertice SaaS Inflation Index, which has run between 12% and 16.4% through 2026, peaked at 14.7% in Q4 2025, and sat at 16.4% in June 2026. Zylo’s 2026 SaaS Management Index, built on more than 40 million licenses and over $75 billion of tracked spend, puts the average enterprise at $55.7 million a year on SaaS, up another 8%, while the application portfolio sits flat at 305. Same number of tools. More money.
We already argued that the seat is finished as a unit of value. That argument is closed. The open question is what the replacement does to your controls, because consumption, resolution and outcome pricing each move a decision that used to belong to procurement into runtime, where nobody is watching.
The Buyer Changed Identity
Jason Lemkin’s line about vendor selection is the one worth sitting with: “more and more, the agents will pick the vendor. And the model is critical to that decision.” An agent choosing between two APIs at execution time is performing a purchase. It reads latency, capability and price, then commits your money.
Every commercial term that matters to that decision has to be machine-readable, or the agent picks on the terms it can parse and ignores the rest. Rate limits, per-unit price, overage behavior, remaining balance. If those live in a PDF signed in March, the agent is buying blind.
The seller’s side of this is blunter. “If you sell per seat, you are the funding source. That is the position you’re negotiating from in 2026.” Per-seat vendors are where the money for consumption-priced AI comes from, which is a procurement fact before it is a pricing opinion.
Resolution Is the Tier You Can Actually Reach
Outcome pricing is the hardest tier to reach. Lemkin’s estimate: “Outcome pricing is hard in roughly 95% of categories.” The attribution problem is why. If a support agent handles a conversation and the customer renews some weeks later, no billing system can honestly claim the renewal.
Resolution sits below outcome and above raw consumption, and it is measurable inside the vendor’s own system. A ticket closed. A question answered without escalation. That measurability is what makes it billable without a quarterly argument.
It also creates an arithmetic problem worth running before you sign. A $2.00 per-conversation rate at a 60% resolution rate is an effective $3.33 per actual resolution. The advertised number and the number you pay differ by 67%, and the difference is entirely determined by a quality metric the vendor controls. Normalize every consumption quote to the unit you actually care about before comparing two of them. We made the same point about comparing agent tiers rather than list prices.
Write the Definition Down Before the Pilot Starts
The governance work here is small and boring, and it happens before the first invoice or not at all. Lemkin’s version: “Write the definition of the billable unit down before the test starts… What counts as resolved? What happens on a follow-up contact about the same issue? Who arbitrates? Get it on paper before the first invoice, not after the first dispute.”
Three questions, one page, signed by both sides. A customer who contacts support twice about the same problem generates one resolution or two, and whichever answer you agree to changes the bill materially. The vendor’s default answer will favor the vendor. That is not bad faith, it is just what an unspecified term does.
The arbitration clause is the one to write down. When the vendor’s telemetry reports one resolution rate and your CSAT data implies a worse one, somebody decides. Name that somebody in the contract.
The Budget Becomes a Runtime Control
Lemkin lists four obligations a consumption-priced vendor owes the buyer. Read them as a control specification, because that is what they are:
- A hard cap. “A hard cap they set, not you. Not an alert. A stop.” An alert is a notification that money already left. A cap is a boundary the system cannot cross.
- Pre-bought blocks with rollover. Committed spend that does not evaporate at quarter end, which removes the incentive to burn budget for its own sake.
- Threshold alerting routed to the buyer. “The person who signs the renewal should find out at 60% of budget, not on the invoice.” The obligation is routing: the admin is not the person who signs the renewal.
- A budget API the agent can query. Lemkin’s own assessment: “Almost nobody has built this. If agents are doing the spending, they need to query remaining budget the same way they query anything else.”
That last one is the piece almost nobody has built, and it is the piece that decides whether autonomous spending is governable at all. An agent that cannot read remaining budget has exactly two states: spending, and stopped by a failure. Give it a balance endpoint and it can degrade deliberately. Switch to the cheaper model as the budget runs down. Batch instead of streaming. Queue non-urgent work until the next block. Ask a human before the last of it goes.
None of that is exotic engineering. It is the same pattern any rate-limited API already implements, applied to money instead of requests. My read on why almost nobody has built it: consumption pricing assumed a human reading a dashboard monthly.
Where the Money Is Coming From
Redpoint surveyed 141 CIOs in March 2026. Of those, 45% said their AI budgets come out of existing software budgets rather than net-new money. Creative Strategies puts a sharper number on the same behavior: only about 28 cents of each incremental AI dollar is net-new IT budget, and the remaining 72 cents comes from something already in the stack.
That view is contested. RBC’s CIO research reaches a different conclusion, finding AI spend is largely new budget. Both cannot be right, and the answer probably varies by company size and sector, so treat 28 cents as a directional claim rather than a settled one. What is harder to dispute is the consequence buyers report directly: 79% of IT leaders hit a price increase at renewal in the past 12 months, 78% got unexpected charges tied to AI features or consumption, and 61% cut planned projects to absorb unplanned SaaS cost increases. Cancelled projects are the real price of unmetered consumption.
The defensive response is already visible. In the same survey, 54% of CIOs are running vendor consolidation programs, only 3% expect AI to lead to more vendors, and 54% said they would rather their incumbent vendor add AI than switch to an AI-native alternative. Incumbency is worth more in 2026 than product superiority, which is a hard thing for a challenger to hear and an easy thing to verify at renewal. We traced the procurement side of this in enterprise AI repricing and in what a pricing page has to say to a machine.
Do This Now
Take one consumption-priced AI vendor already in your stack, and spend an hour on four things.
Find the billable unit definition in the contract. If it is absent, or reads “per successful interaction” with no follow-up rule and no arbiter, that is your first amendment request.
Normalize the price. Divide the per-unit rate by the vendor’s own reported success rate and compare that number, not the list price, against the alternative.
Ask the vendor for two endpoints: remaining balance and hard cap. Not an alert. A stop the vendor’s system enforces, with a value you set. The answer you get tells you how seriously they have thought about autonomous buyers.
Then route the 60% threshold notification to whoever signs the renewal. That one is a configuration change, and it removes a common way a consumption bill becomes a surprise.
Vendors that ship a budget API will win agent-driven selection, because an agent can only commit to spend it can bound. Buyers that ask for one now will get it before the invoice teaches them why they wanted it.
This analysis synthesizes The 3 New Pricing Models in B2B. Pick One, Because The Old One (Just Seats) Really is Dying (SaaStr, Jason Lemkin, August 2026), which reports the Vertice SaaS Inflation Index, the Zylo 2026 SaaS Management Index, the March 2026 Redpoint CIO survey, and the Creative Strategies budget-substitution estimate.
Victorino Group helps engineering and procurement teams turn consumption contracts into runtime controls agents can obey. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation