Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

NVIDIA ships Nemotron 3.5 Lightning, an open 30B MoE built for always-on agents.

Only 3B active params per token, runs on a single consumer GPU, claims 4x throughput and ~30% faster task completion than comparable open models. Live.

Enterprise DNA |
NVIDIA ships Nemotron 3.5 Lightning, an open 30B MoE built for always-on agents.

AI Pulse · AI Trends Pulse

The play

NVIDIA's 30B agent model runs on one consumer GPU, so the cost floor for always-on local agents just dropped hard.

NVIDIA just released Nemotron 3.5 Lightning, a 30 billion parameter mixture-of-experts model designed to run agents that stay awake all day. The clever bit is that it only activates 3 billion parameters per token, which means it fits on a single consumer GPU and claims four times the throughput of comparable open models. Task completion is roughly 30 percent faster, according to NVIDIA’s own numbers.

What matters here is not the architecture. It’s the distribution speed. Within a day of release, Nemotron 3.5 Lightning was live on Ollama, OpenRouter, and Fireworks. That’s the real signal. When a model lands on three major platforms in 24 hours, developers are already building with it before the blog post goes cold. Fast pickup means the model solves a problem people actually have, which in this case is running agents that do not time out or rack up API costs while they wait for the next task.

Why this matters if you run a company

Most agent frameworks today spin up, do a thing, then shut down. That works fine for one-off tasks. It doesn’t work if you want an agent monitoring inbound leads, watching for support tickets, or checking inventory levels all day without burning through tokens. A model that runs locally, stays on, and processes requests quickly changes the economics. You can deploy an agent that listens instead of one that polls. This is the kind of infrastructure we’re building into tools like the Omni Command Centre, where persistent agents handle repeating workflows without constant supervision. If you’ve been waiting for agent infrastructure that doesn’t cost like a SaaS subscription, this is the direction to watch.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.