When Agents Talk to Agents, the Work Goes Dark: Anthropic's Own Warning

TV
Thiago Victorino
6 min read
When Agents Talk to Agents, the Work Goes Dark: Anthropic's Own Warning

Anthropic linked roughly 9,700 survey respondents to their actual Claude usage between April and June 2026, and 93% of those conversations produced an identifiable artifact: code, a document, an analysis, a decision. That is the cleanest first-party read we have on what AI is actually doing inside knowledge work. Buried in the same June 26 report is a sentence that should reorder the priority list for anyone running an AI program: agent-to-agent interactions may become inscrutable to humans. The report that proves measurement is finally getting good is also the report that names the moment measurement starts to break.

The measurement got better, and that is the story

The June Economic Index is not a refresh of the old methodology. Anthropic replaced the prior seven-day sampling window with continuous hourly sampling, and they linked privacy-preserving telemetry to survey responses through CLIO, their automated analysis layer. The result is a dataset where self-reported behavior and observed behavior sit next to each other for the same people. That linkage is rare. Most AI adoption research is either a survey (what people say they do) or telemetry (what the logs show), almost never both, joined at the individual level, at this scale.

What the joined data shows is concrete. Across the linked cohort, 27% report cost savings from their AI use and 68% report learning gains. More than a third, 35% and up, expect AI to handle most or nearly all of their work tasks within twelve months. These are not vendor projections. They are what working people told a researcher while a parallel record of their actual usage was being measured against the claim.

The headline correlation is the one worth sitting with. People who delegate to Claude the most are the most optimistic about their own future labor-market outcomes. Anthropic states it plainly: the heaviest delegators are the most positive about where their careers are going. That cuts against the reflex story where automation breeds anxiety. In this cohort, the people handing off the most work feel the most secure, not the least.

Delegation and optimism point in the same direction

There are two honest ways to read that correlation, and an operator needs both on the table. The optimistic read: delegation is a skill, the people who learn to hand work to an agent see their own leverage increase, and they project that forward as career security. The skeptical read: optimism and delegation may share a cause, such as seniority, role autonomy, or working in a function where AI clearly helps, and the arrow between them is not as clean as the headline suggests.

The June dataset cannot fully resolve which read is correct, and Anthropic does not claim it can. What it can do is make the question answerable over time. With continuous sampling and individual linkage, you can watch whether last quarter’s heavy delegators are still optimistic next quarter, and whether their measured output held up. Correlation today becomes a testable trajectory tomorrow. That is the actual upgrade in this report: instrumentation that turns a finding into something you can track quarter over quarter.

The inscrutability warning is the load-bearing sentence

Here is where the 93% number and the inscrutability warning collide. Today, 93% of measured AI work produces an artifact a human can open and read. The whole Economic Index works because the output of AI labor still lands in a human-legible place: a pull request, a doc, a spreadsheet, a sent message. Measurement is possible precisely because a person is still the endpoint.

Agent-to-agent interaction removes that endpoint. When one agent calls another, negotiates a sub-task, passes structured state, and returns a result, the intermediate steps do not have to surface as an artifact anyone reads. The work happens, the outcome lands, and the reasoning that connected them stays inside a machine-to-machine exchange that no dashboard was built to capture. Anthropic naming this in their own report, while sitting on the best telemetry in the industry, is the tell. The people with the clearest view are flagging where the view goes dark.

This is not a far-future concern. The same report shows a third of workers expecting AI to handle nearly all of their tasks within a year. Tasks handled end-to-end by AI are exactly the tasks most likely to route through agent-to-agent steps. The trajectory the survey measures and the blind spot the warning names are the same trajectory.

Why first-party telemetry is the only honest answer

The instinct, when work goes inscrutable, is to demand more logging. That instinct is right but incomplete. The reason Anthropic can even see the edge of this problem is that their telemetry is first-party: instrumented at the source, continuous, and linked to real outcomes. A company relying on vendor-supplied dashboards and after-the-fact summaries will not see agent-to-agent work go dark. It will simply stop appearing in their numbers, and the numbers will look fine, because what you cannot measure does not show up as a problem. It shows up as a quiet absence.

Better first-party telemetry is what makes the inscrutability problem measurable instead of assumed. If you instrument the boundaries where your agents call other agents, the exchange becomes a record. If you only watch the human-facing endpoints, the exchange becomes a rumor. What separates those two postures is a decision made early: instrumenting the boundary before the work moves somewhere you can no longer see.

Do this now

Map where agent-to-agent steps already exist in your stack, then instrument those boundaries before they carry meaningful work. Find every place where one automated step hands off to another without a human reading the result: an agent that calls a tool that calls another agent, a workflow that chains model calls, a pipeline where the only human checkpoint is the final output. For each one, ask a single question: if this exchange produced a wrong or unsafe intermediate result, would anything in your current logging catch it? If the answer is no, that boundary is already dark, and it is carrying less work today than it will next quarter. Instrument it now, while the volume is low and the patterns are simple. The 93% artifact rate is a current condition, not a guarantee. The companies that keep their AI legible will be the ones that built the telemetry before the work stopped passing through human hands.


This analysis synthesizes the Anthropic Economic Index, June 2026 Report (Anthropic, June 2026). For related reading, see Token Economics Becomes a Board Discipline and Tokenmaxxing and the AI Workforce Inflection.

Victorino Group helps enterprises instrument agent-to-agent work before it goes dark. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation