Enterprise AI is getting an affordability overhaul. IBM and Together AI announced a $240 million multi-year agreement on August 11, 2026 to build the first dedicated, large-scale open-source AI inference cluster on IBM Cloud, powered by NVIDIA’s HGX B300 hardware platform.
The deal signals a direct challenge to the pricing dominance of closed frontier models like GPT-4o and Claude. The goal, according to Together AI CEO Vipul Ved Prakash, is to give enterprise customers “performance comparable to closed frontier models without the high costs of proprietary APIs.”
What the Deal Actually Is
IBM will deploy an NVIDIA HGX B300 cluster inside IBM Cloud infrastructure. Together AI, which runs one of the largest open-source model serving platforms in the world, will operate the inference layer on top. Enterprise customers get access to cost-effective, production-grade inference across open-weight models like Meta’s Llama series, Mistral, and Qwen.
The cluster is expected to become available in the first quarter of 2027. According to NVIDIA, the HGX B300 hardware delivers up to 30 times more AI factory output than prior-generation GPU architectures, which means the economics of inference at scale shift dramatically downward.
The deal follows Together AI’s $800 million Series C round, which closed in July 2026 and valued the company at $8.3 billion. That capital is clearly being deployed fast.
Why Enterprises Should Pay Attention
The enterprise AI market has been dominated by a simple narrative: powerful AI costs a lot, so you pay OpenAI or Anthropic for API access and absorb the token costs. That worked fine for pilot projects. At production scale, it gets painful fast.
A legal firm running document review at scale, a healthcare company processing patient records, or a financial services firm analyzing market data cannot afford to route every query through a premium frontier model. The per-token economics simply do not work.
Open-source models have narrowed the capability gap significantly. Llama 4, Qwen 3.5, and Mistral models now handle a wide range of enterprise tasks without much visible quality difference for structured, deterministic workflows. The bottleneck has shifted from model capability to inference infrastructure: where do you actually run these models reliably, at enterprise scale, with the SLAs your business needs?
That is exactly what IBM and Together AI are solving. IBM Cloud brings the enterprise trust layer, the compliance certifications, and the existing relationships with large organizations already in IBM’s ecosystem. Together AI brings the model expertise, the inference optimization, and the open-source community credibility. NVIDIA’s HGX B300 hardware brings the raw compute to make it cost-effective.
The Pricing Bet
The core thesis here is that enterprises care more about cost per token than they do about which company’s logo is on the model. IBM CEO Arvind Krishna has said publicly that IBM is betting on enterprises prioritising cost and control over prestige. This deal is the infrastructure manifestation of that bet.
For context, Together AI has reportedly achieved inference costs as low as a fraction of what OpenAI charges for equivalent output quality on many tasks. When you multiply that delta across millions of daily queries in a large enterprise, it represents material savings that CFOs actually notice.
The $240 million figure tells you something important: this is not a small experiment. Both IBM and Together AI are committing serious capital to a multi-year vision. They believe enterprise demand for affordable, open-source inference will be large enough to justify building dedicated infrastructure for it.
What This Means for Business
If you are running AI in your business or evaluating where to start, this deal reinforces a strategic direction worth considering: the proprietary versus open-source decision is no longer purely a capability question. It is increasingly an economics question.
For routine, high-volume tasks, open-weight models running on infrastructure like what IBM and Together AI are building will deliver cost structures that closed frontier models simply cannot match. For cutting-edge reasoning tasks, frontier models still have advantages that matter.
The smart approach for most businesses is a hybrid one: use the right model for the right task. Not every query needs the most expensive model. Building that judgment into your AI architecture now, before you are locked into pricing structures designed for single-vendor relationships, is one of the more valuable things you can do.
Enterprise DNA’s Omni platform is built around exactly this kind of architecture thinking. We help businesses identify which AI investments actually pay off, and cost-per-task economics is central to that analysis.
The IBM and Together AI partnership is a signal that the infrastructure layer is maturing in a direction that favours businesses willing to build deliberate AI strategies rather than simply defaulting to whoever has the best marketing.
Ready to think more clearly about your AI infrastructure decisions? Book a session with our team to map out an approach that makes sense for your business.
Source
IBM Newsroom