Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

YC-backed Bullet claims a faster, cheaper coding-agent harness than Claude Code and Codex

Ex-AppLovin/Citadel founders report 95.8% on SWE-bench Verified at 119 seconds per task average, crediting model routing and grep-based search over.

Enterprise DNA |
YC-backed Bullet claims a faster, cheaper coding-agent harness than Claude Code and Codex

AI Pulse · AI Trends Pulse

The play

Treat coding-agent benchmark claims cautiously, test with your repositories, and avoid tools requiring root access without strong independent review.

A new coding agent called Bullet is making a big claim. Built by founders who came out of AppLovin and Citadel, the team says their tool hits 95.8% on SWE-bench Verified, a common test for how well an AI agent can fix real software bugs, and does it in about 119 seconds per task on average. That’s fast compared to Claude Code and Codex, the two names most people think of first in this space. Their explanation for the speed is technical but simple enough: instead of embedding an entire codebase and searching through that, they route tasks to different models and use grep-style search, basically fast text matching, to find what needs fixing.

Before you get excited, know that this claim is getting real pushback. On Hacker News, commenters pointed out that SWE-bench scores across the industry have been creeping up for a while, and some think the benchmark itself may be maxed out or gamed rather than a true measure of skill anymore. Bigger issue for owners: the desktop app reportedly asks for root access on your machine, and there’s no open-source code available to check the claims against. That’s a lot of trust to hand over with nothing to verify it.

Here’s why this matters even if you never touch Bullet. Coding agents are becoming a real category, and vendors are racing to post the flashiest benchmark number to get attention. If you’re evaluating any AI dev tool for your team, ask what access it needs, whether its results are independently checkable, and whether the benchmark it’s citing still means anything. This kind of vetting question is exactly the sort of thing we build into an AI command centre, so you have one place to track tool claims against what your team actually experiences (the original report).

Working With Claude field guide cover

Free Resource

Put what you just read to work

The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.

No spam. Unsubscribe any time.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.