Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

xAI claims Grok 4.5 plus its "Grok Build" agent beat GPT-5.5 and Claude Opus 4.8 on an independent real-work benchmark (GDPval+), with the widest margins in legal and healthcare tasks.

Single sponsor-adjacent benchmark disclosed via Musk's own account, not yet independently replicated. Worth watching, not citing as settled.

Enterprise DNA |
xAI claims Grok 4.5 plus its "Grok Build" agent beat GPT-5.5 and Claude Opus 4.8 on an independent real-work benchmark (GDPval+), with the widest margins in legal and healthcare tasks.

AI Pulse · Frontier Labs Watch

The play

Ignore Grok 4.5 claims until independent benchmarks replicate them, single-source leaderboards are not reliable.

xAI posted a claim this week that its Grok 4.5 model, paired with a tool called Grok Build, outperformed GPT-5.5 and Claude Opus 4.8 on a benchmark called GDPval+. The benchmark is supposed to measure how well models handle real work tasks, and xAI says Grok’s lead was widest in legal and healthcare scenarios. Elon Musk shared the results on his own account, and the original report notes the scores come from a single, sponsor-adjacent test that hasn’t been replicated by independent labs yet.

That last part matters. One benchmark, disclosed by the company building the model, tells you something but not everything. It’s not unusual for a vendor to pick the test that makes their product look best. The fact that this benchmark isn’t widely used or verified means you can’t treat these numbers as settled truth. It’s worth watching, not worth citing in a board deck.

What it means for you

If you’re evaluating models for contract review, patient intake, or other high-stakes workflows, don’t switch based on a single claim. Test the models yourself on your own data. The gap between a controlled benchmark and messy real-world documents is often large. That said, the claim does suggest Grok is pushing hard into domains where accuracy and reasoning depth matter, and competition in those areas benefits everyone who’s building systems that need to handle nuance.

This is exactly the kind of signal we track inside the Omni Command Centre, where you can compare model performance on your own tasks and route work to whichever model actually performs best for you. One vendor’s benchmark is interesting. Your own results are what count.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.