- News
OpenAI Ultrafast: GPT-5.6 Sol Hits 750 Tokens Per Second
OpenAI previews Ultrafast, a new API tier powered by Cerebras that runs GPT-5.6 Sol at 14x standard speed for real-time enterprise use cases.
- News
IBM Backs Together AI With $240M to Cut Enterprise AI Costs
IBM and Together AI announced a $240M multi-year agreement to build a large-scale open-source AI inference cluster on IBM Cloud using NVIDIA HGX B300 hardware.
- News
AMD Acquires Taalas: Chips That Bake AI Models Into Silicon
AMD bought startup Taalas, whose chips encode AI model weights into silicon for 10x+ inference speedup — AI infrastructure is maturing fast.
- News
Fireworks AI Raises $1.5B on Specialized Enterprise AI
Fireworks AI closed a $1.5B Series D at a $17.5B valuation, showing enterprises are moving past generic models toward custom AI built on their own data.
- News
SambaNova Raises $1B to Compete With Nvidia on AI Inference
SambaNova's Series F signals investors are betting big on Nvidia alternatives as JPMorgan and SoftBank back on-premises AI inference at scale.
- Blog
Fireworks AI: What Engineers Actually Found
A practitioner's take on Fireworks AI inference in production, covering latency wins, cost surprises, and where it fits next to Together, vLLM, and OpenAI.
- Blog
Groq Speed: What Engineers Actually Found
Honest look at Groq inference speed in production. Latency numbers, cost surprises, where it works, where it breaks, and what teams pair it with.
- News
OpenAI's Jalapeño Chip Cuts AI Inference Costs by 50%
OpenAI and Broadcom unveiled Jalapeño, a custom inference chip built in nine months that delivers 50% cost savings over standard AI GPUs.
- News
Anthropic Eyes UK Startup Fractile for Inference Chips
The Information reports Anthropic is exploring a fourth chip supplier with Fractile's SRAM-based design that claims 25x faster inference at one-tenth the cost.