Enterprise DNA
News Trending Research

AI Agents That Negotiate: Anthropic's Key Finding

Anthropic's Project Swap ran 205 market experiments with AI agents trading books for 201 employees. The biggest finding has nothing to do with negotiation.

Enterprise DNA | | via Anthropic Research
AI Agents That Negotiate: Anthropic's Key Finding

Anthropic just published the results of one of the most practical agentic AI research projects to date. They called it Project Swap, and on the surface it sounds simple: get 201 employees across six offices to lend their Claude agents to a peer-to-peer book trading marketplace.

But what they found is a direct warning for every business currently deploying AI agents.

What Happened in Project Swap

Each participant brought books they wanted to give away and had a five-minute conversation with Claude about their reading preferences. Their agent then joined a digital trading floor and negotiated exchanges with other employees’ agents.

Anthropic ran 205 separate market experiments, varying the underlying model (Haiku, Sonnet, Opus, and Fable) and the agent instructions (from “maximise your principal’s outcomes” to “be prosocial and help everyone”). Then they surveyed participants on how happy they were with what they received.

The results were illuminating.

The Numbers Worth Paying Attention To

On a 10-book preference list, participants ended up with books ranked around their 5th choice on average. The theoretical optimum — if the market were perfectly efficient — was closer to 2nd.

That’s a meaningful gap. Here’s where it came from:

  • 85% of the shortfall was from imperfect preference elicitation (the agent didn’t fully understand what the person wanted)
  • 15% came from inefficient negotiation between agents

In other words, the agents were actually pretty good at negotiating once they understood the brief. The problem was that a five-minute conversation wasn’t enough to build a complete picture of someone’s tastes.

Preference accuracy hit 61% on average — better than popularity-based recommendations (53%) and collaborative filtering (55%), but still leaving significant gaps. The agents understood you better than an algorithm would, but not as well as you’d want before handing over real decisions.

Model Capability Matters More Than Instructions

This is the finding that every operations and technology leader should write down.

Opus agents achieved a trading efficiency score of 0.88. Haiku agents scored 0.75. That’s a larger gap than anything the researchers found by changing the instructions — whether agents were told to be ruthless or cooperative made far less difference than the underlying model they were running on.

The practical translation: if you’re deploying AI agents for anything consequential — procurement, vendor negotiations, customer escalations, lead qualification — the model you choose has more impact than the prompt strategy you design around it. You can spend months fine-tuning instructions, but upgrading the model will likely do more.

This aligns with what many enterprise teams are discovering in production: prompt engineering has a ceiling, and it’s lower than the ceiling on model capability.

What People Were Willing to Delegate

Sixty percent of participants rated their received books 7.2 out of 10 — not bad for a first-pass autonomous negotiation system. But the trust metric is more telling: participants said they’d be comfortable delegating roughly 30% of their annual book budget to an AI agent.

That’s a number comparable to what people delegate to trusted friends. For a technology that didn’t exist three years ago, it suggests the ceiling on autonomous delegation is higher than most enterprise buyers currently assume.

What This Means for Business

The research has three direct implications for anyone building or buying agentic AI systems:

1. Invest in preference elicitation before you invest in automation. The biggest gains in Project Swap came from helping agents understand what people actually want, not from making the negotiation logic smarter. For enterprise deployments, this means building robust onboarding flows, context gathering, and verification steps before any agent starts acting autonomously.

2. Model choice is a strategic decision, not a cost optimisation. Using a weaker model to cut API costs is reasonable in many contexts. But for high-stakes agentic work — where an agent is taking actions with real consequences — capability differences compound quickly. The research quantifies what many practitioners already suspected.

3. Governance infrastructure matters as much as the agents themselves. Project Swap surfaced a set of requirements for agentic marketplaces: identity verification, dispute resolution, rate limiting, and clear rules about what agents can and can’t do. The same applies to any enterprise workflow. Agents operating without governance infrastructure create downstream liability, not efficiency.

Anthropic plans to extend this research into real economic contexts — moving from book swaps to commercial procurement scenarios. The early results suggest that when AI agents act as market participants, their principals (humans and businesses) need to think carefully about how much those agents actually understand their goals.

The takeaway isn’t that agents can’t be trusted. It’s that trust has to be earned through better preference understanding first, not assumed because the automation is running.


Enterprise DNA’s Omni Ops service helps businesses design agentic workflows that get the preference elicitation and governance infrastructure right from the start. Learn more about Omni Ops or explore how we approach AI agent deployment for business operations.

Working With Claude field guide cover

Free Resource

Going deeper with Claude?

Get the free 32-page implementation guide for ANZ teams.

Add your name (optional)

No spam. Unsubscribe any time.