Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

An independent benchmark puts Claude Code at 2.6x the cost-per-success of the cheapest harness

Composio ran DeepSeek V4 Pro across five harnesses (Claude Code, DeepSeek Harness, Hermes Agent, Pi Agent, OpenCode) on the same 30 tasks with the.

Enterprise DNA |
An independent benchmark puts Claude Code at 2.6x the cost-per-success of the cheapest harness

AI Pulse · Under the Radar

The play

Benchmark harnesses on your own tasks, since cheaper tooling may deliver similar outcomes at materially lower cost.

Most people assume the AI tool wraps around the model do not matter much. This test suggests they matter a lot.

Composio took one model, DeepSeek V4 Pro, and ran it through five different harnesses, the software scaffolding that handles how an AI agent takes actions, checks its work, and retries when something fails. Same model, same 30 tasks, only the harness changed. The result: DeepSeek’s own harness cost $0.028 per successful task. Claude Code cost $0.074 per successful task. That’s about 2.6 times more expensive to get the same job done, with the same underlying model doing the thinking.

Why does this matter if you’re running a business and not a research lab? Because when you’re pricing out AI agents for real work, whether that’s customer support, data cleanup, or research tasks, the model gets most of the attention but the harness is quietly doing a lot of the cost math. Two vendors can offer “the same AI” and hand you wildly different bills, not because the intelligence differs but because one wraps it in leaner, more efficient scaffolding.

A few caveats worth saying plainly. This is one vendor’s test, on their own benchmark, with no independent replication yet. The methodology is disclosed and the model was held constant, which is good practice, but you should treat the exact numbers as a signal to investigate, not gospel.

The practical takeaway is simple: when you’re comparing AI agent tools, ask about the harness, not just the model. This is exactly the kind of tradeoff we track when we build out an AI command centre, since cost-per-success across tools is the kind of thing you want visibility on before you scale usage, not after. You can see the original thread here: the original report.

Working With Claude field guide cover

Free Resource

Put what you just read to work

The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.

No spam. Unsubscribe any time.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.