- Home
- The Thinking Wire
- 77% of Enterprises Re-evaluate AI Vendors Every Six Months. Put Cost per Outcome in the Template
77% of Enterprises Re-evaluate AI Vendors Every Six Months. Put Cost per Outcome in the Template
Madrona surveyed 150 enterprise IT professionals. Per TechCrunch’s write-up of the report, 74% plan to expand their AI budgets in the next 12 months and the rest plan to hold spending steady. Fewer than half of their AI pilots ever reach full production. And 77% of the enterprises re-evaluate their AI vendors every six months or on a rolling basis.
TechCrunch read those numbers as a warning to startups: recurring revenue that renews every six months is a weaker asset than recurring revenue on a multi-year contract. Madrona’s own framing, as quoted, is that this creates “a fast in, fast out dynamic that is fundamentally different from traditional enterprise SaaS, where multi-year contracts provided a moat of inertia.” In enterprise AI, “switching costs are lower and the re-evaluation cadence is relentless.”
Read the same data from the buyer’s chair and it describes something else. Buyer-side AI governance has shown up as a calendar entry. Three quarters of surveyed enterprises already sit down twice a year, or continuously, and ask whether each AI vendor still earns its place. The mechanism exists. What I would add to that meeting is one metric pair that would make it decisive.
Governance arrived as a cadence
Governance programmes usually get written as policy: acceptable-use documents, model inventories, approval workflows. In my experience those take quarters to draft and longer to enforce. A re-evaluation cadence does part of the same work with none of the drafting. Every six months someone has to answer, with budget consequences, whether the vendor delivered.
The two Madrona numbers reinforce each other. Budgets are growing at 74% of the respondents, and 77% of them re-open the vendor decision at least every six months. More spend, reviewed more often. A vendor that survives the review keeps a growing account. A vendor that fails it has no multi-year contract to fall back on.
We covered the investor side of this in ARR broke in the AI era, built on Simon Wu’s interviews. Madrona is a different dataset: 150 enterprise IT professionals, describing their own behaviour. It confirms the pattern from the demand side, which is the side that decides.
What buyers say they want to pay for
TechCrunch also cites a second survey, from Andreessen Horowitz, attributed to Tugce Erten and Sarah Wang. Fifty buyers were asked how they want to pay. More than half want AI fees tied to the work produced or to other outcomes rather than to usage such as the number of tokens consumed.
Fifty respondents is a small sample and the survey has no public link in the article, so treat the direction as the finding, and the exact proportion as indicative. The direction lines up with the Madrona cadence. A buyer who reviews vendors every six months wants a number to review that maps to business results, and token volume does not map to anything a CFO recognises.
Yesterday’s piece, when the LLM adjudicates the invoice, covered outcome pricing from the vendor’s side and the harder problem underneath it: who decides that an outcome was completed. I will not re-argue that here. What matters for the buyer’s template is that the adjudication question has to be settled before cost per outcome can be a review metric at all.
Earlier we tracked the end of flat-fee pricing and the move to per-token billing. Cost per outcome is the next step on the same path. Per-token billing made cost visible per request. Outcome accounting makes it visible per result, and the two do not always move together.
IDC, as cited in the same TechCrunch article, has enterprises on pace to spend $4.25 trillion on technology in 2026. The AI share of that is what the six-month reviews are redistributing.
The metric pair a vendor proposed
The clearest statement of the metric came from a vendor. Steve Sweetman, VP of Product Management for Foundry Models at Microsoft, published a piece on the Azure blog (dated August 26, with the year 2026 in the page metadata) arguing that the number to manage is “the cost of a successful outcome, not the price of a token.”
His reasoning is mechanical. “A single completed outcome can take a dozen model requests.” In agentic workloads the prompt prefix is resent on every turn, so “an agent that takes 10 turns pays for that prefix 10 times.” A team optimising cost per request can lower it and still lose money: “An optimization that lowers the first while raising the number of turns has made things worse, and only the second will show it.”
That is the metric pair. Cost per request and cost per completed outcome, tracked together, over the same period, from your own telemetry. Either alone can mislead. Cost per request drops when a vendor swaps in a cheaper model. If the cheaper model needs more turns to finish, cost per outcome rises and the first metric hides it. Sweetman’s line applies to the buyer as much as the vendor: “You cannot tune what you cannot see, and you cannot claim a saving you did not measure.”
A vendor proposing the metric that would discipline its own pricing is worth taking at face value on this point. It is also the metric that, once in your template, lets you compare vendors that price differently.
Vendor levers, labelled as such
Sweetman’s article lists the levers Azure offers to reduce cost, and every figure in it is a vendor claim with no customer data attached. Batch deployments carry “up to 50% lower costs.” Cache reads are “discounted up to 100% on provisioned deployments.” Both are ceilings, marked “up to”. Model subsets “now align with Azure Policy” and “constrain routing to an approved allow-list where a compliance boundary applies.”
For the review template, the figures matter less than the questions they generate. A vendor that claims batch savings should be able to show your batch share and your realised discount. A vendor that offers caching should be able to show your cache hit rate. And the allow-list item is a governance control regardless of the savings: which models can this product route your requests to, who approved that list, and has it changed since the last review. Routing that silently moves to a cheaper model is a cost-per-request improvement and a compliance event at the same time.
A re-evaluation template
Neither TechCrunch’s write-up nor Sweetman’s post proposes a review template. What follows is my proposal, built from their data. Six items, in the order a review meeting should take them.
- Cost per request and cost per completed outcome, both, over the review period, from your own logs. If the vendor supplies only the first, record that as a finding.
- The definition of a completed outcome, written down, and the name of whoever adjudicates it. If the vendor’s system decides, say so. The adjudication problem is the whole argument in one line.
- Turns per outcome, trending. Sweetman’s dozen-requests figure is a vendor illustration, so measure your own. A rising trend with a falling per-request cost is the failure mode he describes.
- The routing allow-list. Which models, approved by whom, changed when.
- Pilot status. Fewer than half of pilots reach production, per Madrona. The review is where each one is promoted, extended with a written reason, or killed.
- Exit cost, verified. Madrona reports lower switching costs in enterprise AI. Test the claim against your own integration before you rely on it. A vendor you cannot leave is a vendor you cannot review.
Items 1 and 3 are the pair Sweetman proposed. Items 2 and 4 are governance. Items 5 and 6 are the parts of the Madrona data that translate into decisions.
Do this now
Find the next AI vendor review on your calendar. If there is none, your organisation is outside the 77% and the first action is to put one there, six months out at most.
Before the meeting, pull two numbers from your own telemetry for every AI product in scope: cost per request and cost per completed outcome, with your definition of “completed” written next to the second. Then ask each vendor for the same two numbers from their side. A vendor who can answer both has given you a basis for the next six months. A vendor who can answer only the first has told you what their pricing is designed to hide.
This analysis synthesizes Startup ARR Is Less Secure Than Ever, New Research Shows (TechCrunch, Julie Bort, citing Madrona and Andreessen Horowitz surveys, September 2026) and The Economics of Agent Optimization: Four Ways to Lower the Cost (Microsoft, Steve Sweetman, Foundry Models, August 2026). The Madrona report is named but not linked in the TechCrunch text, and its figures are cited here as TechCrunch reported them.
Victorino Group helps engineering and finance leaders build the AI vendor review that measures cost per completed outcome alongside cost per request. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation