AI Pulse · Under the Radar
The play
Test Opus 5 on a non-critical workflow before swapping it into production, it may break existing prompts that assume literal instruction-following.
Claude’s new Opus 5 model landed last week, and the reaction from people actually building with it tells a more useful story than the benchmarks. Dan Shipper at Every says it argued with his instructions, stopped mid-task, and broke existing workflows until his team rebuilt prompts from scratch. Aaron Levie at Box reports the opposite: 19% gains on technology tasks and 13% on healthcare in their internal evals. Anthropic’s own Alex Albert notes it uses fewer tokens at lower reasoning levels than earlier Opus versions.
The pattern that matters is not who’s right. It’s that the model behaves differently. Opus 5 is less literal. It interprets instructions instead of following them word for word, which means it sometimes does what it thinks you want rather than what you said. That autonomy helps in some workflows and breaks others, and right now teams are rewriting prompts and skills to match the new behavior. Reddit’s r/singularity ran five separate benchmark threads in the same window, and the common thread was the same: this is a behavior shift, not just a performance bump.
If you’re running AI tools in production, this is the kind of change that shows up as silent drift. A model update lands, your outputs shift, and you don’t always know why until something breaks or a user flags it. That’s why version control and eval tracking matter more than leaderboard position. When a model starts interpreting instead of executing, you need to see it before your customers do. The discussion on Hacker News has 1,700 upvotes and counting, mostly from builders comparing notes on what broke and what improved.
This is exactly the kind of thing we build into an AI command centre: version tracking, output comparison, and drift detection so you catch behavior changes before they hit production. The model will keep evolving. Your systems need to keep up.
Free daily email
Get this every morning.
This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.
Free daily email
Subscribe to the daily AI Pulse
One short read every morning on what is actually happening in AI. Free.
You are in
Your first AI Pulse lands tomorrow morning. Keep an eye on your inbox.