AI Pulse · AI Trends Pulse
The play
Treat coding-agent benchmark claims cautiously, test with your repositories, and avoid tools requiring root access without strong independent review.
A new coding agent called Bullet is making a big claim. Built by founders who came out of AppLovin and Citadel, the team says their tool hits 95.8% on SWE-bench Verified, a common test for how well an AI agent can fix real software bugs, and does it in about 119 seconds per task on average. That’s fast compared to Claude Code and Codex, the two names most people think of first in this space. Their explanation for the speed is technical but simple enough: instead of embedding an entire codebase and searching through that, they route tasks to different models and use grep-style search, basically fast text matching, to find what needs fixing.
Before you get excited, know that this claim is getting real pushback. On Hacker News, commenters pointed out that SWE-bench scores across the industry have been creeping up for a while, and some think the benchmark itself may be maxed out or gamed rather than a true measure of skill anymore. Bigger issue for owners: the desktop app reportedly asks for root access on your machine, and there’s no open-source code available to check the claims against. That’s a lot of trust to hand over with nothing to verify it.
Here’s why this matters even if you never touch Bullet. Coding agents are becoming a real category, and vendors are racing to post the flashiest benchmark number to get attention. If you’re evaluating any AI dev tool for your team, ask what access it needs, whether its results are independently checkable, and whether the benchmark it’s citing still means anything. This kind of vetting question is exactly the sort of thing we build into an AI command centre, so you have one place to track tool claims against what your team actually experiences (the original report).
Free Resource
Put what you just read to work
The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideWant this working inside your business?
See what's possibleFree daily email
Get this every morning.
This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.
Free daily email
Subscribe to the daily AI Pulse
One short read every morning on what is actually happening in AI. Free.
You are in
Your first AI Pulse lands tomorrow morning. Keep an eye on your inbox.