The Lawyers Are Running the Evals

TV
Thiago Victorino
8 min read
The Lawyers Are Running the Evals

An NDA came back redlined in four minutes and fifty-three seconds. A lawyer working the same document with no technology took around 58 minutes. The person reporting that number attaches his own hedge to it: “a 10x lift, though in a sterile environment: a known sample, one contract type.” That caveat is the honest part, and it is not the interesting part.

The interesting part is who tuned the judges that decided the redline was good.

Crosby is a law firm, with the obligations that word carries. It employs lawyers, carries malpractice liability, and bills fixed fees for delivered outputs rather than hours. It has raised $85M from Sequoia, Lux, Index, Elad Gil and BCV, and counts Ramp, Cognition, Cursor and Clay among its customers. Kirkland & Ellis, on the incumbent side of the same market, is spending $500M building internal technology. Both facts describe the same pressure. Only one of them describes a change in who owns the verification loop.

The Problem Nobody Can Buy Their Way Out Of

Coding agents got good fast because code is verifiable. A test suite runs, a compiler complains, a benchmark scores. The reward signal is cheap, mechanical, and available in unlimited quantity. Every serious agent product in software rides on that free verifier.

Contract review has no compiler. Whether a limitation-of-liability clause is acceptable depends on the counterparty, the deal size, the client’s risk appetite, and what got negotiated in the last three agreements with that same vendor. Two competent lawyers will disagree, and both will be defensible. There is no oracle to call.

That leaves exactly two options for anyone deploying agents in a domain like this: pretend the problem does not exist and ship on vibes, or build the verifier yourself. Crosby built it. Sharan Ramjee, their founding research engineer, states the destination plainly: “the question we’re driving toward is how to turn contract review into a more verifiable domain by building custom reward models. Eventually, this can pave the way to autonomous review through reinforcement learning.”

That sentence is a governance statement before it is a research roadmap. If you intend to run agents autonomously in a domain where correctness is contested, you first have to manufacture the thing that says what correct means. And then somebody has to own it.

Who Holds the Pen on the Judge

The loop Crosby runs has four moving parts. Online evals detect regression on every generated edit, so quality drift shows up in production rather than in a quarterly review. An offline dataset auto-funnels failure cases into a growing regression suite. An LLM actively discovers new failure modes instead of waiting for a human to notice them. And the lawyers tune the judges.

That last one is the load-bearing piece. The people who write the eval criteria are the same people who answer to the client when the redline is wrong. Ariba Khan, a member of their technical staff, describes the posture: “You need online evals, continuously learning from them… you can never assume that you’re at the frontier.”

Compare that to how most regulated functions are adopting agents right now. A platform team or a vendor builds the evaluation harness. Compliance reviews the output sample once at procurement. The practitioners who hold the professional liability, the lawyers, the auditors, the clinicians, the risk officers, receive a system whose definition of “acceptable” was written by people who will never be sued for it.

That arrangement fails in a specific and predictable way. The judge encodes what the builder could measure, not what the professional is accountable for. The metrics stay green while the failures that matter accumulate underneath, because the failures that matter were never in the rubric. Nobody notices until a client does.

Fee Structure Is a Control

The billable hour rewards time spent. An agent that compresses 58 minutes into five destroys revenue under that model, which is why incumbent firms have to route AI gains into capacity or margin games rather than passing them through. Crosby sells fixed-fee outputs. Speed converts directly into throughput and margin, so there is no organizational incentive to slow the work down or pad it.

The fee structure is what lets their eval loop stay honest. When shaving minutes helps the firm, the measurement of whether output quality held becomes a genuine question rather than a threat. Under billable hours, an eval that says “the agent is fast and accurate” is an eval that argues against the person who commissioned it.

Look at your own incentive structure before you trust your own quality metrics. If a function’s budget or headcount shrinks when the agent performs well, the people running the evals are being asked to grade their own replacement. They will grade it kindly or harshly, but they will not grade it neutrally.

Measurement as Ritual, Not Report

Crosby runs timed internal time trials: the same document raced with and without the harness, scored on a public contract review scoreboard, tracked week over week. A recurring, visible, competitive event, rather than a benchmark run once at launch.

That design does two things a quarterly quality report cannot. It keeps the baseline alive, so “we got faster” remains a measured claim rather than an anniversary anecdote. And it makes the practitioners participants in the measurement rather than subjects of it. A lawyer who raced the harness and lost knows precisely which part of the output he had to fix, which is exactly the input the eval judge needs next.

There is also a product decision here that functions as policy. As the reporting puts it, “you can either make the lawyers feel like they are at the service of a robot or they are in control, directing them and orchestrating.” Interface design determines whether the person with the liability behaves like an operator or an approver. An approver rubber-stamps. An operator catches things.

Ross Weiser, a lawyer who left Sullivan & Cromwell for legal engineering at Crosby, describes what the clients are actually buying: “They want to know that they can trust us. They want to know who’s responsible. It is just so human.” Clients are not procuring throughput. They are procuring a named party who owns the outcome. Every governance control worth building traces back to that question.

Do This Now

Take one agent-assisted workflow in a regulated or judgment-heavy function of your company. Contract review, credit decisions, clinical documentation, marketing claims, whatever you have running. Then answer three things in writing.

Who wrote the eval criteria, by name and role? If the answer is a platform team, a vendor, or nobody, the person holding the liability is not holding the verifier. Move the pen.

When did the eval last change, and what triggered the change? If failure cases are not auto-funneling into a regression set, your quality signal is a snapshot from the day you launched.

Does the accountable function lose budget, headcount, or billable revenue when the agent performs well? If yes, your quality metrics are compromised at the source, and no amount of eval sophistication fixes that. Fix the incentive first.

We have written before about legal AI at vendor scale and about traceability shipped as a product feature. Crosby adds the piece those miss. In a domain with no compiler, the verifier is a governance artifact, and it belongs to whoever answers the phone when the work is wrong.


This analysis synthesizes Crosby: The Inner Workings of a Neofirm (New Ontologies, 2026).

Victorino Group helps regulated functions design eval loops owned by the people who carry the liability. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation