Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News Trending Product

Google Releases Gemini 3.6 Flash and Confirms Gemini 4

Google launched three new Gemini models on July 21, cutting agent token costs up to 71%, while confirming Gemini 4 pre-training has begun.

Enterprise DNA | | via 9to5Google
Google Releases Gemini 3.6 Flash and Confirms Gemini 4

After three missed deadlines on Gemini 3.5 Pro, Google shipped something else entirely. On July 21, the company released three new Gemini models — Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber — and in the same announcement confirmed that Gemini 4 pre-training has begun, the first time Google has publicly overlapped a shipping model with a confirmed next-generation flagship in the pipeline.

What Launched

Gemini 3.6 Flash is the headline model. Priced at $1.50 per million input tokens and $7.50 per million output tokens, it costs less per token than the previous Gemini 3.5 Flash at $9 per million output tokens. The real saving goes deeper than the sticker price: the model produces roughly 17% fewer output tokens to complete the same tasks, which puts the effective per-task cost reduction at around 31% compared to 3.5 Flash. For agentic coding workloads specifically, that figure jumps to around 71%.

Benchmark results reflect the improvement. On DeepSWE — a coding agent benchmark — 3.6 Flash scores 49%, up from 37% for 3.5 Flash. On MLE Bench it moved from 49.7% to 63.9%. On OSWorld-Verified, which tests a model’s ability to control real software interfaces, it improved from 78.4% to 83%.

Gemini 3.5 Flash-Lite targets the high-volume, cost-sensitive end of the market at $0.30 per million input tokens and $2.50 per million output tokens. This positions it as Google’s answer to lightweight classification, filtering, and routing tasks where running a frontier model would be overkill.

Gemini 3.5 Flash Cyber is a security-tuned variant of the 3.5 Flash model, aimed at teams building threat detection, compliance automation, and secure document analysis workflows.

All three are available now through the Gemini API, Google AI Studio, Android Studio, and the Gemini Enterprise Agent Platform.

Gemini 4 Is Coming

Alongside the Flash launches, Google confirmed it has begun pre-training Gemini 4, describing it as its most ambitious pre-training run to date. No timeline was given, but the public confirmation — while shipping a cheaper, faster model — signals a deliberate two-speed strategy: the Flash tier competes on price and throughput today, while Gemini 4 is being positioned as the benchmark-leader for enterprise trust once it ships.

Gemini 3.5 Pro, which missed its original June target at Google I/O and two subsequent July deadlines, remains unavailable. Google has not given a revised release date.

What This Means for Business

Token cost is not a line item most businesses think about until they start running AI agents at scale. Once you have agents handling multi-step tasks across real workflows — pulling data, drafting responses, updating records, routing decisions — the cost per token multiplies fast. What looks cheap in a demo becomes a real budget line in production.

The Gemini 3.6 Flash numbers change that math. A 31% reduction in effective per-task cost means an agentic deployment that was costing $10,000 a month could cost $6,900 instead, with no change in output quality. For agentic coding specifically, the economics improve even more.

The three-tier release also gives businesses more options for architectural decisions. You can use 3.5 Flash-Lite for fast, cheap routing and classification — work out which category a customer query falls into, or whether a document needs escalation — and deploy 3.6 Flash for the tasks that actually require reasoning. That’s a smarter architecture than running your most expensive model on every step.

The security-tuned Cyber variant is worth watching for teams building on top of Google’s infrastructure who handle regulated or sensitive data. Compliance-sensitive industries — finance, healthcare, legal — often need to demonstrate that the AI systems they use have been configured with security in mind, not just bolted on afterwards.

On Gemini 4: the fact that Google is publicly confirming pre-training while shipping discount models is unusual. It suggests the company is trying to hold enterprise attention through a period where it has been slower than competitors on frontier model releases. Businesses evaluating Google as their long-term AI platform can treat the pre-training announcement as a signal that the next generation is real and coming, even if the timeline remains unclear.

The practical takeaway for businesses running or planning to run AI agent workflows: Google has just made its API infrastructure meaningfully cheaper. If you’re building on Google’s stack, the price drop is real and the benchmarks support the efficiency claims. If you’re on another provider’s stack, this release is a competitive reference point worth using in your next vendor conversation.

Working With Claude field guide cover

Free Resource

Going deeper with Claude?

Get the free 32-page implementation guide for ANZ teams.

No spam. Unsubscribe any time.