AI Pulse · AI Trends Pulse
The play
Track inference costs and latency by workload, while keeping GPU providers flexible as dedicated decode hardware enters agent infrastructure.
Nvidia says its Groq 3 LPX inference chip is now in full production. It comes from Nvidia’s $20 billion Groq acquisition and is designed for the part of AI work that happens after a model has been trained, when it is generating responses token by token. Nvidia reports 500MB of on-chip memory, 256 chips per rack on its Vera Rubin platform, and a benchmark result of 3,400 output tokens per second on a Gemma 4 31B test with a 100,000-token context window. Nebius is the first cloud provider deploying it, according to Nvidia’s announcement.
For a business owner, the key point is not the chip specification. It’s that AI providers are building separate infrastructure for AI agents that need to read, reason and respond quickly over long workflows. GPUs remain essential, particularly for training and broad AI workloads. But Nvidia is making a clear structural bet that high-volume agent responses, often called decoding, need dedicated hardware too.
That matters because faster and more predictable inference can make AI practical in customer service, internal operations and research workflows where waiting several seconds for each step breaks the experience. It may also help providers manage their own compute bottlenecks as demand rises. This is the kind of thing we build into an AI command centre, tracking the infrastructure shifts that will affect the cost, speed and reliability of AI inside your company.
Free Resource
Put what you just read to work
The free 32-page Working With Claude guide: the full ecosystem, Claude Code, and how to roll it out across a business.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideWant this working inside your business?
See what's possibleFree daily email
Get this every morning.
This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.
Free daily email
Subscribe to the daily AI Pulse
One short read every morning on what is actually happening in AI. Free.
You are in
Your first AI Pulse lands tomorrow morning. Keep an eye on your inbox.