- Home
- The Thinking Wire
- In Banking, McKinsey Puts Human Oversight at 70 to 75 Percent of an Agent's Variable Cost
In Banking, McKinsey Puts Human Oversight at 70 to 75 Percent of an Agent's Variable Cost
For an agent doing a customer service task in banking, token costs frequently land at 20 to 25 percent of the variable run cost. Human oversight accounts for 70 to 75 percent. That split comes from McKinsey QuantumBlack, published August 24, 2026, and it puts a price tag on the part of an agent program that almost never appears in a cost review.
Model routing, prompt caching, rightsizing to a cheaper tier: the optimization moves McKinsey says organizations reach for first all attack the 20 to 25. The 70 to 75 is set somewhere else entirely, by how many runs a human has to look at, how long each look takes, and who is qualified to take it. Those are governance parameters. They now have a dollar figure attached.
The Number Behind the Number
The oversight line is a product of one variable McKinsey names directly: the exception rate. In banking customer onboarding, they would expect 10 to 20 percent of agentic runs to be reviewed by risk and functional experts. A completed bank-account-opening workflow, in their example, involves five to seven agents, multiple deterministic systems, and two to four teams of people providing oversight.
Two to four teams. That is the cost object. Not a shared services queue and not an afterthought in the launch checklist, but standing capacity from expensive people, sized by a review rule written during design, and worth asking when it was last revisited.
The economics still work in the customer’s favor. McKinsey estimates that total cost to complete the onboarding workflow can fall from roughly $50 to $150 per customer to roughly $10 to $30, using standard benchmarks. The agent earns its place. What the breakdown reveals is where the remaining cost concentrates once the agent has already won the comparison, and that is a different question from whether to deploy.
Scale behaves differently on the two lines, too. Using a conversational agent to onboard 2,500 new customers per year costs $10,000 to $15,000 by McKinsey’s estimate; doubling the volume of new customers increased costs to just $15,000 to $20,000. Cost per run compresses as volume grows, in McKinsey’s own numbers. Review capacity does not compress the same way, because a review is a human hour and the exception rate applies to the larger denominator.
Model Choice Is the Visible Lever, Which Is Why It Gets Pulled
Pricing makes the token line easy to reason about. As of early July 2026, McKinsey cites a frontier model such as OpenAI’s GPT-5.5 at about $5 per million input tokens and $30 per million output tokens, against $0.20 and $1.25 for an earlier lightweight model. Roughly 25x on input and about 24x on output, taking those figures at face value. Those ratios are legible on a slide, they map to a procurement decision, and someone can own the savings.
McKinsey identifies three factors driving the token cost of a run: quality of model selected, latency required, and task frequency. All three sit with the engineering team. All three move the smaller number.
Their own recommendation is unusually direct about where that leads. “Many organizations focus their optimization efforts on model selection and token costs because those are the most visible expenses. The more productive approach is to redesign workflows to reduce exception rates and simplify review processes that require costly human interventions.”
Visibility, then, is the selection mechanism. The token bill arrives monthly from a vendor with a line-item breakdown. The oversight bill arrives as headcount inside a risk function, allocated to a cost center that has nothing to do with the agent, and no one reconciles the two.
T(verify) Now Has a Price
We wrote in July about the agent deployment inequality: deploy when P(success) exceeds T(verify) divided by T(do). That framing treats verification time as a ratio, which is exactly right for the deploy decision and silent about what happens afterwards. An agent can clear the inequality easily and still carry a review burden that dominates its operating cost for years.
This breakdown supplies the missing term. T(verify) multiplied by the exception rate, multiplied by run volume, multiplied by the loaded cost of whoever is qualified to review, is 70 to 75 percent of what the agent costs to run in McKinsey’s banking example. Lari Hämäläinen, the McKinsey senior partner who appeared in that July interview, is among this article’s authors.
Three design choices set that product, and none of them are model choices.
The routing rule decides which runs a human sees. A rule that sends every run above a confidence threshold to review will behave very differently in cost from one that samples, or one that reviews by transaction value, or one that reviews only where a write is irreversible.
The escalation path decides who reviews. Routing an exception to a risk expert instead of a first-line operator can change the hourly cost of that review by a large multiple, and an escalation design will default upward whenever that feels safer at design time than being wrong.
The exception rate itself is a workflow property. Cases that break because an upstream system returns an ambiguous value, or because the agent was handed a task with underspecified acceptance criteria, generate reviews that a workflow change would eliminate.
What This Costs If You Ignore It
McKinsey’s own range for what a bank can spend is wide: customer-facing agents can cost $20,000 to $30,000 to run a single-agent workflow, and $100,000 to $200,000 to run a multiagent team, based on their analysis of public research and public pricing information. McKinsey does not break that range into fixed and variable, so the split cannot be applied to it directly. What the split does say is that within variable cost, oversight is the dominant line. Because oversight runs at roughly three times the token line there, a change that cuts the exception rate materially moves more money than a change that only cuts token price.
Treat those numbers as McKinsey estimates for illustrative banking cases, because that is what they are. The structure of the argument survives the uncertainty in the figures. Whatever your split turns out to be, you cannot manage it without measuring it, and in the agent programs I have looked at, review hours per run are not instrumented at all.
The fixed costs deserve a mention because they compound the same way. McKinsey splits them into AI infrastructure (public cloud containers, memory, management, analytics) and agent orchestration, which they define as the data scientist capacity required to maintain and enhance the agent in production. They note production agents often need to be tweaked every couple of days as new foundation models, model context protocols, enterprise systems, and business requirements emerge. Both fixed lines are people lines as well.
Do This Now
Instrument review hours per agent run before your next model-selection debate. You need three fields per run: whether it was reviewed, who reviewed it, and how long the review took. Agent logging that captures tokens and latency and stops there leaves the token line as the only one anyone can argue about with data. It is the same instrumentation problem we described in cost per completed task, one layer further out.
Then put the exception rate on the same page as the model choice in whatever forum approves agent spend. McKinsey recommends an AgentOps capability analogous to FinOps, and putting AI unit economics into quarterly business reviews. The specific artifact matters less than the adjacency: a review routing rule that costs six figures a year belongs in the same conversation as a model tier that costs five.
The teams that get this right will look like they found a cheaper model. They will have redesigned an escalation path.
This analysis synthesizes Where AI agents pay off: A practical guide to the economics of agentic workflows (McKinsey QuantumBlack, August 2026), Is that AI agent worth it? Agentic economics and the modern operating model (McKinsey Quarterly, July 2026).
Victorino Group helps engineering organizations instrument review cost per agent run and redesign the routing rules that drive it. Let’s talk.
All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →
If this resonates, let's talk
We help companies implement AI without losing control.
Schedule a Conversation