Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News Trending Product

Gemini 3.7 Flash: 50% Off, Bigger Gains on Agent Benchmarks

Google's Gemini 3.7 Flash drops coding costs by 50% and posts major benchmark gains on agentic tasks, arriving just three weeks after 3.6 Flash.

Enterprise DNA | | via VentureBeat
Gemini 3.7 Flash: 50% Off, Bigger Gains on Agent Benchmarks

Google has launched Gemini 3.7 Flash, a new AI model tuned for coding and autonomous agent workflows, at an introductory price that cuts the cost of running AI agents roughly in half compared to its predecessor. The release arrives on August 13, 2026, just three weeks after Gemini 3.6 Flash — a turnaround Google attributes to developer feedback and targeted improvements to the model’s reasoning core rather than a full retraining run.

For businesses running AI at any scale, this is worth paying attention to. Lower inference costs mean the economics of agentic AI just got a lot more attractive.

What Changed Under the Hood

Gemini 3.7 Flash is not a ground-up rebuild. Google made algorithmic improvements to the reasoning foundation of 3.6 Flash and added customizable “thinking configurations” that let developers tune the balance between output quality, cost, and latency depending on the task.

The practical result is a model that applies more disciplined multi-step planning and produces fewer tool errors during complex agent runs. Early enterprise adopters report the model achieves prompt-cache hit rates around 8% higher than 3.6 Flash, which compounds into meaningful cost savings on high-volume deployments.

Benchmark Gains That Matter for Agents

The numbers on coding and agentic benchmarks are substantial:

  • DeepSWE v1.1 (software engineering tasks): 65.3% vs 48.6% on 3.6 Flash
  • FrontierCode 1.1 Main (production-grade code generation): 43.6% vs 34.4%
  • AutomationBench (multi-step automation workflows): 30.4% vs 17.0%
  • GDP.pdf (enterprise document comprehension): 34.0% vs 22.0%

These are not marginal improvements. A model that solves 65% of real software engineering tasks versus 49% is a qualitatively different tool for any business deploying AI coding agents or workflow automation.

On competitive comparisons, Gemini 3.7 Flash scores 85.8% on Terminal-bench 2.1, sitting just behind GPT-5.6 Terra’s 87.4%. Claude Sonnet 5 leads on multimodal desktop task benchmarks with 33.3% versus 3.7 Flash’s 26.3%. The model is most competitive in the cost-efficiency bracket, not the raw capability tier.

The Pricing Story

Through December 31, 2026, Gemini 3.7 Flash costs $0.75 per million input tokens and $3.75 per million output tokens. That is half the rate of Gemini 3.6 Flash at launch.

Starting January 1, 2027, prices rise to $1.50 per million input and $7.50 per million output — still competitive, but the year-end window is the headline. Businesses running high-volume agent workflows who move quickly could lock in meaningful savings before the pricing steps up.

When you factor in the reduced tool errors and higher cache hit rates, enterprise deployments of Gemini 3.7 Flash are reporting effective costs roughly 35% lower than running the same workloads on 3.6 Flash. At scale, that math changes what’s viable to automate.

What This Means for Business

The release tells two stories simultaneously.

Story one is cost normalisation. The price of AI inference continues to fall faster than most planning assumptions from six months ago. Teams that built business cases around 3.6 Flash’s pricing can now run the same workloads for less, freeing budget to expand scope or run more agent tasks in parallel. The companies moving fastest on AI adoption are benefiting from compounding efficiency gains that weren’t in any 2025 forecast.

Story two is the cadence itself. Three weeks between significant model releases is a pace most enterprise IT procurement cycles can’t match. Businesses that built rigid integrations around a specific model version are discovering those integrations age quickly. The smarter approach is building agent workflows on abstraction layers that let you swap models without rewriting everything.

For teams evaluating AI agent infrastructure right now, Gemini 3.7 Flash is worth a serious benchmark run, particularly for coding-intensive pipelines, document processing, and multi-step workflow automation where the DeepSWE and AutomationBench gains translate directly to business outcomes.

Google is making Gemini 3.7 Flash available immediately through the Gemini API, Google AI Studio, Android Studio, and the Gemini Enterprise Agent Platform. Vertex AI customers can access it through Google Cloud’s enterprise tier.

The pace of improvement means whatever you’re planning to build with AI agents — the cost and capability envelope you’re working within today looks materially different to six months ago. The businesses acting on that are the ones pulling ahead.

Working With Claude field guide cover

Free Resource

Going deeper with Claude?

Get the free 32-page implementation guide for ANZ teams.

No spam. Unsubscribe any time.