Your Corrections Are the Judgment You Failed to Front-Load

TV
Thiago Victorino
7 min read
Your Corrections Are the Judgment You Failed to Front-Load

Melody Koh of NextView Ventures spent a month doing something almost nobody running agents does: she kept score on herself. Across fifteen of her own work sessions, she logged every time she interrupted an agent and sorted each interruption into one of three buckets. Was it irreducible judgment, something only she could decide? Was it mechanical overhead, a click or a paste the setup should have handled? Or was it correcting a mistake the agent made?

Corrections won. By a wide margin, in her words. She publishes no percentage, and the honest reading is that the exact split matters less than the ranking.

Her conclusion in The Autonomous Middle is uncomfortable in a useful way. “The correction load was the price of the judgment I’d failed to front-load.” Those interruptions trace back to her own unstated preferences, arriving late and one turn at a time, in the form of “no, not like that.”

We have argued before that verification is the real bottleneck and that judgment goes unmeasured. This piece is the instrument for the second claim. It is a small audit you can run on your team’s last fifteen sessions this week, and it produces a number nobody currently has.

The Three Buckets

Pick a session. Walk its transcript. Every point where a human typed something that was not part of the original task definition is an intervention. Classify each one.

Irreducible judgment. A decision that required your taste, your context, or your accountability. Which of two viable directions to take. Whether a claim is defensible enough to publish. Whether the tradeoff is acceptable to the customer. This bucket is the work. It never goes to zero and should not.

Mechanical overhead. Copying output from one tool into another. Re-running a command the agent could have run. Fetching a file the agent had no path to. Approving the same class of action for the eleventh time. This bucket is a tooling defect wearing the costume of oversight.

Correcting agent mistakes. The agent produced something wrong, and you caught it. Koh’s examples are specific and familiar: “the agent would invent an attribution, or talk itself out of a problem it had already flagged.” A fabricated source. A concern raised in one paragraph and rationalized away in the next.

The classification is harder than it sounds, and the difficulty is where the value sits. Most teams call all three “review.” Bundled together they produce one useless conclusion, which is that agents still need supervision. Separated, they point at three different owners: the human’s, the platform’s, and the prompt’s.

Run It on Fifteen Sessions

The audit needs no instrumentation to start. A spreadsheet with four columns works: session, intervention, bucket, and one line of what you actually said.

Fifteen sessions is roughly a month of real work for one person, and it is the number Koh used. Do not sample only your worst sessions; the ones that went smoothly carry the information about what front-loading already works. Log the intervention at the moment it happens if you can, because reconstruction after the fact quietly reclassifies corrections as judgment. Nobody remembers being wrong as often as they were.

One rule keeps the data honest. When you are unsure whether an intervention was judgment or a correction, ask a single question: could I have written this down before the session started? If yes, it was a correction. The test covers everything that was writable in advance, whether or not anyone wrote it.

What the Buckets Mean When Corrections Dominate

A correction-heavy profile is the common result, and it reads as good news once you understand what it says. Corrections are the bucket most under your control, because each one is a rule you already hold and have not yet externalized.

The response is to move the correction into standing form. Koh’s practice is to make every correction a rule that gets enforced automatically on the next run rather than remembered. Her review agents read each draft cold through fixed lenses, one for accuracy, one for voice, one for whether the piece lands. A correction caught by a human once becomes a check applied by a reviewer every time.

Overhead-heavy is a platform bill. If your interventions cluster around moving data between systems and approving repetitive actions, no amount of prompt engineering fixes it. Someone needs to own the connective tissue.

Judgment-heavy, with the other two small, is the mature profile and the rarest. It means the routine has been separated out and what remains genuinely requires you. If a team reports this on the first audit, check the classification before celebrating. Judgment-heavy on paper is usually correction-heavy with generous labelling.

The Sandwich Relocates the Tax

The structural move Koh describes is one she credits to Kieran Klaassen of Every: the AI sandwich, human judgment at the front, agent work in the middle, human judgment at the exit. Klaassen’s compound engineering framing is his, not hers. Ethan Mollick’s “software-brained” is likewise borrowed.

Koh is precise about what the sandwich buys, and the precision is the part worth carrying into your own operation: “The sandwich doesn’t pay that tax down; it relocates it, out of the middle where it stalls you turn by turn, to the edges where it doesn’t.” Review is a fixed cost. A team that expects front-loading to reduce total review time will read its own metrics as failure. The win is that the same total review no longer arrives as interruptions that break the agent’s run and your attention at once.

Her original contribution sits underneath that: the width of the autonomous middle is a trainable skill, not a property of the model you happen to be using. “The limit on autonomy was never really the model. It’s how cleanly you can separate the judgment from the routine.” Two people running identical tooling will get different middles. The one who front-loads better gets the wider one.

The Residue You Cannot Front-Load

Koh estimates that the last ten to twenty percent of judgment is unforeseeable and only surfaces at exit review. That figure is an estimate she offers, not a measurement, and it should be treated as a shape rather than a target.

The shape matters for how you staff the exit. If a fifth of the judgment on a piece of work is undiscoverable until the work exists, then exit review cannot be delegated to whoever is free. It has to be the person with the accountability, and it has to be scheduled as real work rather than squeezed into the twenty minutes before a meeting. Teams that front-load well and then rush the exit have optimized the cheap half of the sandwich.

Do This Now

Take your last fifteen agent sessions. Log every intervention into one of the three buckets, using the writable-in-advance test to separate judgment from correction. Then take the single most frequent correction and convert it into a rule that runs without you, either as a check in the setup or as a lens a review agent applies to every output.

One correction converted per week is fifty-two standing rules a year, each one bought with evidence rather than guessed at during a retrospective. The audit is small enough to run on a Friday afternoon and it produces the only honest answer to a question most engineering leaders are currently answering by feel: how much of what we call oversight was avoidable?

Note the caveat that Koh states herself. Fifteen sessions, self-observed and self-classified, is a disciplined self-audit and not a study. Treat her numbers as a method to copy, not a benchmark to hit. Your fifteen sessions will produce your own ranking, and yours is the one that tells you where to spend the next month.


This analysis synthesizes The Autonomous Middle (NextView Ventures, July 2026).

Victorino Group helps engineering organizations instrument the judgment work that agent metrics leave invisible. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation