- Home
- The Thinking Wire
- Outcome Pricing Needs a Judge. Zendesk's Judge Is a Language Model.
Outcome Pricing Needs a Judge. Zendesk's Judge Is a Language Model.
Zendesk bills only for what it calls Verified Resolutions, confirmed by an LLM evaluation within 72 hours of the conversation. Assisted escalations and contained resolutions cost nothing. The billable rate sits at roughly $1.20 to $1.50 on committed volume.
On a pricing page that reads as a buyer win. As a control document it reads differently: a language model is deciding whether a payment obligation exists.
The Shape of the Market
Three vendors, three different answers to the same question, and they are worth separating carefully because the labels blur.
Intercom charges $0.99 for each conversation its Fin agent resolves, and nothing at all for the ones it does not. Zendesk charges for Verified Resolutions only, with the LLM check described above. Salesforce launched at $2 per conversation, charged for every 24-hour session whether or not anything was resolved, then replaced that with Flex Credits at about 10 cents each, starting at $500 for 100,000 credits.
The Salesforce move is the one worth reading slowly. Flex Credits are consumption pricing. You buy units and you spend them. That is closer to the metered cloud model everyone already understands than it is to paying only when the work succeeded. The original $2-per-session was not outcome pricing either, since it charged whether or not anything got resolved. Salesforce went from time-boxed metering to credit metering. Neither version makes the vendor’s revenue contingent on the result.
Keith Kirkpatrick, research director for enterprise software at Futurum Group, puts it plainly: “Outcome-based pricing is becoming a market standard.” The direction of travel is clear. What the direction implies operationally is what I have not seen worked through.
And then there is OpenAI. Per Kevin McLaughlin and Amir Efrati at The Information, OpenAI has begun letting some of its largest customers pay only when its AI actually completes the job. The arrangement is limited to select major accounts rather than offered generally. The terms, the customers and the prices are all unknown. The Information sits behind a paywall and is the originating report here; The Next Web, which relayed it, states that it has not independently verified it. Everything checkable in this essay comes from the Intercom, Zendesk and Salesforce mechanics, which are published. OpenAI is a signal that the pattern is reaching the largest accounts, and nothing more than that.
Somebody Has to Rule on Success
The structural consequence never appears on the pricing page. The moment payment depends on whether the AI succeeded, success stops being a description and becomes a determination. A determination requires a determiner.
For a support conversation the candidates are limited. The customer could rule, by rating the interaction, which introduces mood, response bias and a long tail of people who never answer. A human reviewer could rule, which is accurate and destroys the unit economics that made outcome pricing attractive in the first place. A deterministic rule could rule, something like “no human touched this ticket within seven days”, which is cheap and auditable and wrong often enough that both sides would argue about it.
Or a model could rule. That is the option Zendesk shipped. Within 72 hours of the conversation, an LLM evaluation confirms whether a resolution happened. Its verdict creates or does not create a line on the invoice.
Of the vendors described in the reporting, that LLM evaluation is the only adjudication mechanism anyone has actually specified. The reporting says what a resolved conversation costs at Intercom, not who determines that it was resolved. OpenAI’s terms are unknown. So the honest reading is that one vendor has documented its judge and the rest have not said.
The Questions Procurement Cannot Yet Ask
Financial controls assume a billing input is either a measurement or a contract term. Meter readings, seat counts, API calls, GB-months. These are contestable in a boring way: you disagree, you pull the log, you reconcile.
An LLM verdict is neither a measurement nor a term. It is an inference. Inferences have error rates, and this one is attached to money.
Four questions follow, and I have not seen any vendor answer them in public:
- Who audits the grader? The party that benefits from a resolution being billable is the party that owns the model that decides it was billable. Bad faith is not the concern. The concern is an unaudited incentive sitting inside a billing input, which I would expect any procurement review to flag if it appeared anywhere else in the contract.
- What is the false-negative rate? No accuracy rate is published for the Zendesk grader, in either direction. The false-negative rate matters to the vendor’s revenue. The false-positive rate matters to the buyer’s invoice. Nobody outside the vendor can currently size either.
- Can the buyer see the evaluation transcript? For a disputed meter reading, the log is the evidence. For a disputed model verdict, the evidence would be the evaluation itself: what the model saw, what it concluded, on what basis. No source I read describes a mechanism for a customer to request it.
- What is the dispute process when the model is wrong? The model will be wrong sometimes. Any error rate above zero says so, and no vendor claims otherwise. A pricing page that establishes a payment trigger without establishing an appeal path has assigned the risk of that error entirely to one side.
I am posing these as open questions on purpose. No source publishes an accuracy rate, an appeal path or a transcript policy for any of these graders. The absence of published answers is the finding. Do not read it as evidence that the answers are bad.
The Billing Unit Moved Again
We argued earlier that the billing unit is the control surface: whatever you are charged for is where governance actually has to live, because that is the thing the vendor optimizes and the thing your finance team can see. Outcome pricing does not weaken that argument. It relocates it somewhere much stranger.
When the unit was a token, governing spend meant governing usage, which is why cost variance beats spending caps as a control. When the unit was a task the agent chose to run, the question became who approves the agent’s budget. When the unit is a verified outcome, the control surface is a classifier. You cannot cap it, you cannot rate-limit it, and you mostly cannot see it. You can only ask how it was validated, which is a model-governance question wearing a purchase order.
That vocabulary already exists in your organization. Model risk management, challenger models, validation reports, error analysis on a labelled sample. Model risk management has run this playbook on credit models for years, which is where I would go looking for the language. I have not seen any of it reach a software procurement conversation, because the vendor’s classifier has not previously sat between the work and the invoice.
Ask for the Grader Spec Before You Sign
If you have an outcome-priced AI contract on the table this quarter, add one section to the requirements before legal starts marking it up. Call it the adjudication spec, and require the vendor to state in writing:
- What determines a billable outcome, in enough mechanical detail that you could describe it to your CFO without using the vendor’s marketing terms.
- Whether that determination is made by a model, a rule, a human, or the end customer. If it is a model, whether the vendor will state its measured accuracy and on what evaluation set.
- Whether you can obtain the per-item evaluation record for any charge you dispute, and within what window.
- The dispute and correction path, including who pays while a disputed charge is being reviewed.
- Your right to sample. A quarterly right to pull a sample of billed items and review them yourself is the single cheapest control here, and it is the one most likely to be granted, because it costs the vendor nothing if the grader is good.
A vendor that answers all five has built something defensible and should be happy to say so. A vendor that cannot answer the second one is asking you to accept an unaudited model as an accounts-payable trigger. That is a reasonable thing to decline, or at minimum a reasonable thing to price into the deal.
Outcome pricing genuinely is better aligned than seats or tokens. That alignment is real, and it is why the model is spreading. It just arrives carrying a governance obligation that no line on any published pricing page has yet acknowledged, and the buyer is the party that inherits it.
This analysis synthesizes OpenAI has started letting some customers pay only when the AI works (The Next Web, Ana Maria Constantin, August 2026) and the originating report, OpenAI Starts Letting Customers Pay When AI Works (The Information, Kevin McLaughlin and Amir Efrati, August 2026), which is paywalled and was not read directly.
Victorino Group helps engineering and procurement teams write the adjudication controls into AI contracts before the invoice arrives. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation