Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News Trending Product

OpenAI Ultrafast: GPT-5.6 Sol Hits 750 Tokens Per Second

OpenAI previews Ultrafast, a new API tier powered by Cerebras that runs GPT-5.6 Sol at 14x standard speed for real-time enterprise use cases.

Enterprise DNA | | via Cerebras Systems
OpenAI Ultrafast: GPT-5.6 Sol Hits 750 Tokens Per Second

OpenAI announced on August 13 that it is previewing a new API tier called Ultrafast, which runs GPT-5.6 Sol at up to 750 output tokens per second — roughly 14 times faster than standard processing. The speed comes from a partnership with Cerebras Systems, whose wafer-scale chips keep model weights entirely on-chip rather than streaming them from external memory.

The launch matters for any business running AI at the point of a customer conversation, a live incident, or a fast-moving data workflow. At 750 tokens per second, an AI agent can complete a multi-paragraph analysis in the time it previously took to finish a sentence.

How Cerebras Makes This Possible

Standard GPU-based inference has a bottleneck that most enterprise buyers never see but always feel: the model weights live in off-chip memory and have to be loaded repeatedly during generation. As models get larger, this memory bandwidth problem gets worse.

Cerebras’ Wafer-Scale Engine keeps 44 GB of SRAM directly on the chip itself, eliminating the round-trip. For a frontier model like GPT-5.6 Sol, this translates to a 5.6x end-to-end speedup on GDP-Val, a benchmark measuring economically valuable knowledge work tasks, with no measurable drop in output quality.

The result is that Ultrafast delivers the same intelligence as the standard GPT-5.6 Sol tier, just at a pace that makes genuinely real-time applications viable for the first time at frontier model quality.

What OpenAI Is Targeting

OpenAI is positioning Ultrafast squarely at time-sensitive enterprise workflows. The use cases they are emphasising in the early preview:

  • Incident response and debugging: when systems go down, engineers need answers in seconds, not the time it takes for a slow model to reason through logs
  • Financial research and fraud detection: markets move faster than standard inference can keep up with for real-time signal analysis
  • Real-time customer support and voice: voice AI has always been constrained by latency; 750 tokens per second changes what is conversationally possible
  • Commerce and live experimentation: personalisation decisions made mid-session require inference that matches the pace of a click

The early access cohort includes Jane Street for financial research, Podium for commerce applications, and Rogo for financial analysis. That group tells you something about where OpenAI thinks the immediate demand is concentrated: financial services and any business where the AI has to respond before the human’s patience runs out.

Why Voice AI Teams Should Pay Attention

Voice applications are the most latency-sensitive use case in the enterprise AI stack. A delay of even a few hundred milliseconds makes a voice AI employee sound hesitant or broken in conversation. Standard inference speeds have been the limiting factor that forced voice AI developers to choose between model quality and conversational responsiveness.

Ultrafast dissolves that tradeoff. At 750 tokens per second, a voice AI employee can generate a full response before the audio pipeline even needs it, which means the constraint shifts from inference speed to network latency, a much smaller and more predictable variable.

For businesses considering AI employees for customer-facing roles, inbound support, or internal knowledge retrieval, this is the kind of infrastructure shift that makes the product category meaningfully better.

Availability and What Comes Next

Ultrafast is launching as a limited preview through the OpenAI API. Access is being expanded gradually as Cerebras scales capacity. OpenAI has not announced public pricing for the tier, but the positioning alongside Jane Street and Podium suggests it is being structured as a premium API option for latency-critical production applications, not a toy or a developer experiment.

The limitation worth noting: availability is restricted. Businesses that want access now need to apply rather than flip a switch.

What This Means for Business

The pattern of this announcement is familiar: a speed improvement that looks like a developer metric turns out to be a business capability unlock. Ultrafast follows that pattern.

Real-time AI has been constrained by inference latency since the category began. That constraint shaped the kinds of applications businesses could build, the kinds of workflows agents could join, and the user experiences that felt natural versus robotic. Removing 14x of that friction does not just make existing applications run faster. It opens use cases that were never viable before.

For businesses running AI in any real-time context, whether that is voice, live customer interactions, incident response, or financial workflows, Ultrafast represents a shift worth watching closely. The companies in the early cohort are treating it as a capability unlock, not just a cost optimisation.

Enterprise DNA works with businesses deploying AI employees and agent workforces across these exact use cases. If real-time inference speed has been a constraint in a deployment you are evaluating, this development changes the calculus. Talk to us about what Omni Voice or Omni Ops looks like at Ultrafast speeds.

Working With Claude field guide cover

Free Resource

Going deeper with Claude?

Get the free 32-page implementation guide for ANZ teams.

No spam. Unsubscribe any time.