Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

OpenAI and Cerebras launch a 750 tokens/second inference tier for GPT-5.6 Sol

"Ultrafast Mode" runs GPT-5.6 Sol on Cerebras wafer-scale hardware, claiming roughly 7x the speed of comparable models on Humanity's Last Exam and a.

Enterprise DNA |
OpenAI and Cerebras launch a 750 tokens/second inference tier for GPT-5.6 Sol

AI Pulse · Frontier Labs Watch

The play

For latency-sensitive workflows, watch Cerebras access and test whether faster responses improve completion rates enough to justify premium infrastructure.

OpenAI just rolled out a new speed tier for GPT-5.6 Sol, and it’s a big enough jump that it’s worth pausing on. Called “Ultrafast Mode,” it runs the model on Cerebras wafer-scale chips instead of the usual Nvidia setup, and OpenAI is claiming about 7x the speed of comparable models on a tough reasoning benchmark called Humanity’s Last Exam, plus a 5.6x speedup on GDP-Val, which tests knowledge-work tasks closer to what your team actually does day to day. Right now access is limited to a select group of customers, so this isn’t something you can just switch on yet, according to the original report.

Here’s why it matters even if you’re not on that early list. Speed at this level changes what AI is useful for inside a business. When a model responds in a second or two instead of ten or twenty, you can start using it for things that need to feel instant: live customer chat, real time document review while someone’s on a call, or an assistant that checks numbers as fast as someone can type a question. Slow AI gets used for batch work. Fast AI starts getting used in the middle of actual conversations and decisions, which is a different kind of tool entirely.

The bigger business story is that OpenAI is spreading its bets across hardware providers instead of relying only on Nvidia. That’s a supply chain decision, not a feature, but it matters for reliability and pricing down the line if you’re planning to build serious workflows on top of these models. This kind of raw speed is exactly the sort of upgrade that becomes more valuable once it’s wired into a real system rather than used one prompt at a time, which is the kind of thing we build into an AI command centre at Omni.

Working With Claude field guide cover

Free Resource

Put what you just read to work

The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.

No spam. Unsubscribe any time.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.