Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News Trending AI News

DeepSeek-V4-Flash-0731 dropped and is dominating HN and Hugging Face trending.

Open-weights (MIT license), 284B total/13B active params, 1M context, priced at $0.14/1M input and $0.28/1M output tokens with a $0.003/1M cache-hit.

Enterprise DNA |
DeepSeek-V4-Flash-0731 dropped and is dominating HN and Hugging Face trending.

AI Pulse · AI Trends Pulse

The play

DeepSeek V4 Flash cuts inference cost by half versus GPT-4o at comparable quality, run a side-by-side on your highest-volume use case this week.

DeepSeek just released V4-Flash-0731, and the numbers explain why it’s sitting at the top of Hacker News and trending hard on Hugging Face. This is an open-weights model under an MIT license, meaning you can run it yourself or fork it without permission. It uses 284 billion total parameters but only activates 13 billion at a time, which keeps inference cheap and fast. Context window is one million tokens.

The pricing is where this gets interesting for anyone running real workloads: $0.14 per million input tokens, $0.28 per million output tokens, and a $0.003 per million cache-hit rate. That last number matters if you’re processing similar queries or documents repeatedly, because cached tokens cost almost nothing. Artificial Analysis scored it 50 on their Intelligence Index, ranking it third out of 101 evaluated models. That puts it in the same performance tier as closed models that cost multiples more per token.

This is not hype. The cost-to-performance ratio is a genuine shift, especially for businesses that need high-volume inference or want to self-host without licensing restrictions. If you’re spending thousands a month on API calls to closed models, this is worth testing against your actual use cases. If you’re building tools that need to scale cheaply, open weights at this performance level change the math.

The kind of decision you face now, comparing this to Claude or GPT-4 variants for specific tasks, is exactly what we build into the Omni Command Centre: a single place to test prompts across models, track cost per task, and route workloads to the cheapest option that hits your quality bar. You don’t need to rebuild your stack every time a new model drops. You need a layer that lets you swap models in and measure the difference in dollars and output quality.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.