The People Who Buy Agents Are Writing the Agent Audit Standard

TV
Thiago Victorino
7 min read
The People Who Buy Agents Are Writing the Agent Audit Standard

AIUC-1, the agent certification standard AIUC announced alongside a $40M Series A, reported by TechCrunch on September 15, 2026, has its test list shaped by a consortium of roughly 250 security and risk leaders who meet monthly. According to the company’s account relayed by TechCrunch, those people are explicitly the buyers of agents. The test list originates on the procurement side of the table.

That single sourcing decision is the most consequential thing about the company, and it is easy to miss under the funding number. The round was led by Ribbit Capital, with First Harmonic participating, on top of a prior $15M seed from Nat Friedman’s NFDG, Emergence, Terrain and Anthropic co-founder Ben Mann. Total funding is $55M. Ribbit is a fintech investor, which is a fact about the cap table and nothing more. What AIUC sells, as TechCrunch describes it, is a standard and a testing service.

What the thing actually is

The design lineage is stated plainly in the TechCrunch report: “Using the widely adopted cybersecurity standard SOC 2 as its muse, AIUC has developed a standard called AIUC-1 and a testing service to validate agents against the standard.”

SOC 2 is a useful anchor because most engineering leaders have lived inside it. It is not a law and not a guarantee. It is a published control framework plus an audit performed by an outside party, producing a report that a buyer reads before signing a contract. The report does not promise the vendor is safe. It says an independent auditor examined specific controls and wrote down what they found. Procurement teams have learned to read that document, argue about its scope and treat its exceptions as negotiating material.

AIUC-1 copies that structure onto agents. The company says an audit runs about 5,000 tests, organised around three risk classes it names: jailbreaks, hallucinations and data leaks. The output is a report of roughly 100 pages per audited agent, detailing where the agent performs safely and where it does not. Those figures come from AIUC and were relayed by a TechCrunch journalist. Nobody outside the company has verified them.

The operating model is worth stating exactly as the company describes it: AI agents run the tests, AI analyses the results, and humans verify the final audit. That one sentence is all the operating model we get. It does not say how many humans, what they check, or what happens when the machine analysis and the human verification disagree.

Named customers so far: Cursor, Lovable, Harvey, ElevenLabs.

Who authors the test list decides what gets tested

A standard is a list of questions. Whoever writes the list controls what counts as risk, and everything absent from the list is, operationally, not a risk at all.

Safety research produces threat models from first principles, asking what a system could do under adversarial pressure. The output is rigorous and frequently unbounded. It tends to include failures that have never occurred in a commercial deployment, because the point of research is to anticipate them.

A consortium of roughly 250 security and risk leaders meeting monthly produces a different list. What a group like that brings, in my experience of procurement committees, is the failure that embarrassed them last quarter, the question their own board asked, the clause their legal team could not get comfortable with. My expectation is that the list comes out narrower than a researcher’s and better calibrated to what has already gone wrong in production. Neither property is measured in the reporting.

Both lists are legitimate. They are not the same list, and the difference shows up in what an audit report can tell a buyer. An agent that passes roughly 5,000 tests built around jailbreaks, hallucinations and data leaks has a report saying so. Whether that suite tracks the failures experienced practitioners have actually been burned by is precisely what nobody outside the company can check yet. A report like that is still worth real money in a procurement conversation. It is not the same claim as “this agent is safe,” and the company’s own framing does not make that claim either.

We have argued before that governance became a procurement decision. This is the next move in that direction, and it is a sharper one. Procurement is no longer just the buyer of governance. Procurement is the author of the specification.

The independence question is structural, not personal

Rajiv Dattani, an AIUC co-founder, was METR’s COO from 2024 to 2025 and remains a board member there. METR does similar testing for the frontier labs, focused until recently on performance rather than safety. Dario Amodei published an essay that floated requiring frontier labs to use embedded third-party evaluators, and named METR as one possibility.

Lay that out as a shape rather than an allegation. One organization tests the agents that companies buy. A second tests the labs that build the underlying models. A person sits on the board of the second and co-founded the first. If third-party evaluation of AI becomes a regulatory expectation rather than a commercial nicety, that arrangement will get examined, and it should be, because the value of an independent audit is entirely a function of how independent the auditor is. Nothing in the reporting suggests impropriety. The overlap is simply a fact about how small this field still is, and small fields have a habit of accumulating governance obligations faster than they accumulate people.

What AIUC-1 has not yet demonstrated

Nothing in the reporting indicates AIUC-1 has been independently reviewed, or shows that the standard discriminates between a safe agent and an unsafe one, which is the only property that matters in an audit regime. A test suite can be large and still miss the thing that breaks you. SOC 2 earned its authority slowly, as buyers compared reports, found the scope games and tightened what they would accept. AIUC-1 is at the beginning of that process, with $55M raised and four named customers.

The honest read: this is a well-capitalized attempt at a structure that has worked before in a different domain, with an authorship model that is genuinely interesting and an evidence base that is currently a single TechCrunch article containing the company’s own numbers.

The third path

Until now, a company buying an agent had two options for handling the risk. Absorb the liability and litigate over the output, which is the position Google was left in when a German court ruled on AI output liability, and the shape of how statutory exposure attaches to agent-generated copy. Or take the vendor’s word for it, backed by a marketing page and a model card.

An independent audit read before signing is a third option, and it is the one that scales. It moves the evidence burden onto the vendor, and it produces an artifact a risk committee can file and compare across vendors. That comparison pressure is what eventually makes a standard mean something.

Do this now

If you buy agents: pull your last three agent vendor contracts and find the clause where the vendor asserts safety. Ask what evidence backs it. If the answer is a model card and a support email, you have a governance deficit you can quantify today, and you now have a vocabulary for what to demand instead. Start asking for third-party test results with named scope. You do not need to wait for AIUC-1 specifically to become the standard.

If you sell agents: assume a buyer-authored test list is coming for your product, whether it is this standard or another one. The work of making your agent auditable is the same work either way. Instrument it so someone outside your company can measure jailbreak resistance, hallucination rate and data containment, then run those measurements yourself before a customer does. Related: reliability in regulated AI comes from the harness.

Either way, the question to put on your next architecture review agenda is short. If an outside auditor asked us to prove this agent behaves, what would we hand them?


This analysis synthesizes Early Anthropic hire, former METR COO have found a way to rein in rogue AI agents (TechCrunch, September 2026).

Victorino Group helps companies make their agents auditable before a buyer or a regulator asks. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation