Claude agents traded 201 people's books. Reading their tastes was the hard part
Anthropic's book-swap experiment suggests an AI negotiator may bargain capably yet still misunderstand the person it represents.
OddBrief EditorialAI-assisted, human-reviewed
AIKey facts
- Participants
- 201 Anthropic employees across six offices
- Preference match
- 61% of ranked book pairs
- Market outcome
- 0.55 against a theoretical 0.89 optimum
- Published
- September 24, 2026
Anthropic published a September 24 experiment in which Claude-powered agents bargained over books for 201 employees. The striking result was not that agents could trade: it was that most of the gap between a good deal and the best possible deal came from misunderstanding what their human owners wanted.
A miniature market with real books
Employees at six Anthropic offices brought books they were willing to give away. Each person briefly described their reading tastes to Claude, which assembled a ranking of books and sent an agent to negotiate with other agents on a digital trading floor. Participants also privately ranked a smaller set of books so researchers could compare the outcome with their actual preferences.
This was a constrained barter market, not an online store. Every participant started with one book, and the agents could propose bilateral swaps or multi-person exchanges. Deals required the relevant agents to agree. The researchers also reran simulated floors with different models and instructions, giving them a way to separate the quality of the bargain from the quality of the information supplied to each negotiator.
Anthropic says Claude's ranking, inferred from the intake conversation, agreed with a participant's own ordering on 61% of book pairs. Random ordering would score 50%. That is a meaningful signal from a short chat, but it also leaves substantial room for the agent to be confidently wrong about the person it is meant to serve.
The negotiators inherited bad maps
The researchers scored the books people received against their own rankings. The decentralized market averaged 0.55 on a scale where 1 means a participant got their top choice. Given the available books and competing preferences, the study calculated a best possible assignment of 0.89. A market run using the agents' inferred preferences, but scored against the human rankings, reached only about 0.60.
Anthropic attributes roughly 85% of the shortfall to imperfect preference information and the remaining 15% to bargaining on the trading floor. That distinction matters for companies proposing shopping agents. Improving negotiation tactics will do little if the assistant has the wrong brief. The study also found that stronger models improved trading outcomes when assessed against the agents' own rankings, though model choice could not erase errors in what the agents thought people wanted.
Participants' book preferences are easier to elicit and score than the constraints in a home purchase, job search, or medical decision. Anthropic's workers also knew they were in an experiment. The findings are therefore a useful warning about delegation, not evidence that autonomous agents are ready to make high-stakes transactions.
A question before the checkout button
The paper suggests a practical design priority: an agent should expose and test its understanding of a person's preferences before it starts dealing. A five-minute conversation can supply a starting point, but the experiment shows how much information remains missing even when the subsequent trading process works fairly well.
Anthropic reported that most recipients liked their books, and participants said on average they would hand Claude about a third of their annual book budget for similar choices. That is a measure of stated willingness in this small setting, not a commercial adoption forecast. The next test is whether real users will correct an agent's assumptions when money, privacy, and irreversible commitments are involved.
Sources
- Project Swap: What happens when agents trade for us?Anthropicprimary source


