Back in June, OpenAI promised its Jalapeño chip would cut AI inference costs by around 50%. On August 25, at the Hot Chips conference at Stanford University, the company published the first real benchmarks. The numbers back up the claim — and for businesses building agentic AI workflows, they go further than expected.
Tested against Nvidia’s GB200 and GB300 systems across three publicly available models — GPT-OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T — Jalapeño delivered between 1.5x and 1.9x more AI work per watt at peak throughput, and between 1.7x and 3.6x lower end-to-end latency. The benchmark methodology used InferenceX, a public standard developed by research firm SemiAnalysis that simulates the full lifecycle of serving an AI request.
The results were measured using OpenAI’s own self-reported data, and some context is important: Jalapeño uses newer HBM4 memory, which gives it an advantage not present in standard Nvidia Blackwell configurations. Nvidia’s upcoming Rubin platform, which also uses HBM4, would be a more like-for-like comparison. OpenAI also focused purely on inference — training workloads remain Nvidia’s domain.
Even with those caveats, the agentic workload numbers are notable.
The Agentic Edge
For AI workflows where many inference steps chain together — agents that reason, retrieve, plan, and act in sequence — Jalapeño showed 2.1x to 4.1x better performance than Nvidia’s comparison hardware. This is the performance range that matters most for businesses building AI that does real multi-step work.
This matters because agentic applications are fundamentally more compute-intensive than single-prompt interactions. An AI agent processing a customer query, checking a CRM, drafting a response, and logging the outcome involves multiple inference calls that compound. At scale, inference efficiency determines whether an agentic deployment is economically viable or a cost sink.
Jalapeño was designed from the ground up around how large language models actually serve requests — optimizing memory movement, networking, and compute patterns for inference specifically, rather than adapting a general-purpose GPU. The 700W power draw (versus 1,400W for Nvidia’s flagship) is a direct consequence of that design philosophy.
Deployment Timeline
OpenAI plans to begin deploying Jalapeño in its own data centers later this year. Volume production is expected to ramp through 2027. For now, external access remains indirect — businesses using OpenAI’s API benefit as the company brings more Jalapeño capacity online and potentially adjusts pricing to reflect lower underlying costs.
This is not a product you can buy. It’s infrastructure that changes the economics of OpenAI’s own services, with downstream effects on pricing and availability for enterprise customers.
What This Means for Business
The cost curve is moving in your favor. OpenAI presented these benchmarks the same week Nvidia reported earnings — a deliberate signal that the era of exclusive GPU dominance over AI compute is being actively contested. More competition in AI infrastructure eventually means lower prices for everyone building on top of it.
Agentic AI just got more compelling. The most significant number from Hot Chips is that 2-4x performance advantage on agentic workloads. If you’ve been building cases for deploying AI agents across operations and the compute cost math hasn’t fully worked, the trajectory is moving your direction. Use cases that looked marginal a year ago are moving toward viable.
Don’t wait for the perfect moment. Businesses that start building AI capability now — even on current infrastructure — will be better positioned to benefit when cheaper compute arrives. The skills, processes, and integrations take time to develop. Getting that foundation right while costs come down is a better strategy than waiting until everything is cheap and then starting from scratch.
The Nvidia relationship changes, not disappears. Custom silicon covers inference. Training frontier models remains an Nvidia-dominated domain. Businesses don’t need to take a position on this infrastructure competition — but understanding that AI inference costs are heading down is useful context for any AI investment decision you make in the next 12 months.
The technical details at Hot Chips are upstream from most business decisions, but the signal matters: AI infrastructure is maturing fast, the costs of running AI at scale are falling, and the case for deploying real AI agents in production workflows gets easier to make every quarter.
If you want to work through where AI fits in your operations — and how to build the right capability now while costs continue to fall — book a session with Sam McKay. Or if your team wants to build the data and AI literacy to evaluate these decisions internally, Enterprise DNA’s learning platform gives you that foundation.
Source
TechCrunch