Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

Google DeepMind ships two efficient Gemini tiers while the flagship 3.5 Pro stays delayed.

3.6 Flash and 3.5 Flash-Lite/Cyber cut token usage up to 17% for agentic workloads, DeepMind's product lead says 3.5 Pro is still in partner testing.

Enterprise DNA |
Google DeepMind ships two efficient Gemini tiers while the flagship 3.5 Pro stays delayed.

AI Pulse · Frontier Labs Watch

The play

If agent task volume is driving your token bill up, test the new Gemini Flash tiers this month for cost relief before budgets lock.

Google just released two new Gemini models, 3.6 Flash and 3.5 Flash-Lite/Cyber, and both are built for efficiency rather than raw power. The headline number is a 17% cut in token usage for agentic workloads, the kind where an AI assistant runs multiple steps or calls tools repeatedly. That matters because tokens are how you pay for API usage, and agentic tasks burn through them fast.

Meanwhile, the flagship 3.5 Pro is still stuck in partner testing with no public launch date. DeepMind’s product lead framed it as a cost and throughput improvement, not a leap in capability. That’s telling. The timing lines up with reports that Chinese AI models now account for roughly 45% of enterprise token use in the US, a shift driven by price and speed more than anything else. Google is clearly playing defense on operating cost, not trying to wow anyone with a new benchmark score.

What this means for you

If you’re running agents or automation that leans on a language model, these Flash updates could trim your bill without changing much else. The real question is whether you’re locked into one provider or set up to switch when economics shift. Chinese models are gaining share because they’re cheaper and fast enough for most business tasks, and now Google is responding by making its mid-tier models leaner.

This is the kind of model selection and cost tracking we build into the Omni Command Centre, so you can see which model fits which task and what it actually costs you per workflow. The days of picking one flagship model and calling it done are over. You need a system that lets you route work to the most efficient option without rewriting code every time a new tier drops.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.