Your Agent Trades Well. It Doesn't Know What You Want.

TV
Thiago Victorino
7 min read
Your Agent Trades Well. It Doesn't Know What You Want.

Anthropic let Claude agents trade books on behalf of 201 of its employees, spread across six local pools, and scored the result against each person’s own ranking of the books. The best possible assignment would have reached 0.89. People ended at 0.55.

The distance between those two numbers is the interesting part of Project Swap, because Anthropic split it in two. In its words, “working from Claude’s imprecise rankings accounts for a majority (85%) of the shortfall, and sending agents into a ‘free-for-all’ trading floor accounts for the remaining 15%.”

Most of the lost value traced back to what the agents believed before they traded. The agents bargained on a picture of their person that was only a little better than a guess.

The picture was barely better than chance

Anthropic compared Claude’s ordering of books with each person’s own. “Claude’s ordering agreed with theirs 61% of the time (where random guessing would achieve 50%).” A baseline that simply ranked books by popularity landed at about 53%.

That puts the agent about eight points above popularity and about eleven above a coin. For a system about to make exchanges in someone’s name, that is a thin margin. Every later decision inherits it.

The trading floor was the easy part

Measured against the agents’ own picture of preferences, the trading was close to optimal. “Haiku trading floors averaged 0.75, while on Opus floors they averaged 0.88 (the utilitarian optimum on Claude’s rankings is 0.95).”

Put next to the 85/15 split, the shape of the problem gets clear. Given what the agents believed, the bigger model traded near the ceiling. The loss sat upstream, in what they believed.

This is also where the title of this piece comes from, and it deserves a hedge. “Trades well” rests on two numbers only: the Opus floors reaching 0.88 against a 0.95 optimum on Claude’s rankings, and the trading floor accounting for 15% of the shortfall. Both are measured on the agents’ view of their people, in a market for books.

Model choice and tone, measured on the agent’s own picture

Anthropic reran the market with variations. Two results are easy to over-read.

“Upgrading an agent from Haiku to Opus moved people 0.12 up their lists.” And “an agent told to be ruthless scored about 0.02 higher than one told to be prosocial.” Both come from reruns scored on Claude’s rankings, not from the live event.

Footnote 15 adds a general limit: “even a large design difference on Claude’s list shrinks to almost nothing on people’s own lists.” A design choice that looks meaningful inside the agent’s model of the person mostly evaporates when you score against what the person actually wanted.

The ruthless-versus-prosocial finding has its own limit. The instructions were two fixed texts. Anthropic’s footnote cites Imas et al., where custom instructions explained much of the outcome difference. What the study shows is narrower: these two instructions, measured this way, barely moved the result.

For a team choosing where to spend effort, the practical reading is uncomfortable. In this study, the lever that mattered most was how accurately the agent knew whom it represented.

Honest agents can still miss the target

Lying was rare. “Only about 1 in 100 agents who mentioned their top pick lied about it. Instructions made no difference here.” Where agents did state their person’s top pick, they almost always stated it honestly. The weak point was the belief, not the reporting of it.

On trust, “the average answer was about 30%,” against about 40% for a well-read friend.

We covered a neighbouring question in agents choosing your vendors, and a control on a sales agent’s discount authority in the discount guardrail piece. Project Swap adds the step before both: does the agent know what its principal wants?

People who act for others pass a test first

Anthropic draws the comparison itself: “Human agents have to pass tests before they can act for others.” It cites the Series 65 and Series 7 exams and FINRA’s know-your-customer rule. My reading of those two: one checks what the client needs, the other checks competence before the adviser may act.

An agent deployment can have the equivalent of both: a check of the agent’s picture of its principal, run before the agent is allowed to act.

What the study can and cannot support

The caveats are real, and they bound everything above.

Anthropic ran the study on its own models and its own employees, who it says are “not representative of the general population.” The stakes were low: books. Only well-behaved agents were tested. The final survey response rate was about 60%. The 0.12 and 0.02 figures are rerun results on Claude’s rankings, and the instructions were two fixed texts.

None of that undoes the central split. It does mean you should treat 85/15 as one well-documented data point about one market, and test your own case before borrowing the ratio.

My reading, which is an inference from this single study: for any agent that buys, sells or negotiates on someone’s behalf, the step the data supports most is intake before delegation. Check what the agent believes its principal wants, then let it act.

Do this now

Build an intake test into every agent that transacts for a person or a business unit.

  1. Before delegation, have the principal rank a small sample of options the agent will face. Keep it short enough that people actually finish it.
  2. Have the agent produce its own ranking of the same sample, from whatever context it has.
  3. Compare the two with a pairwise agreement score. Project Swap gives you reference points: 50% is chance, about 53% is what popularity alone produced, 61% is what Claude reached.
  4. If agreement sits near the popularity line, the agent has learned what is popular and little about the person. Collect more preference data before it acts, or keep a human on the decisions.
  5. Rerun the test when the principal’s situation changes, and log the score next to every delegation so an audit can see what the agent knew when it traded.

Spend the evaluation budget here before spending it on a model upgrade or a tone prompt. In Project Swap, that is where most of the missing value was.


This analysis synthesizes Project Swap: What happens when agents trade for us? (Anthropic, September 2026).

Victorino Group helps teams design intake and delegation controls for agents that act on someone’s behalf. Let’s talk.

All articles on The Thinking Wire are written with the assistance of Anthropic's Opus LLM. Each piece goes through multi-agent research to verify facts and surface contradictions, followed by human review and approval before publication. If you find any inaccurate information or wish to contact our editorial team, please reach out at editorial@victorinollc.com . About The Thinking Wire →

If this resonates, let's talk

We help companies implement AI without losing control.

Schedule a Conversation