Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News AI News

Extreme local inference of frontier-scale open models keeps escalating.

"Run Kimi K3 using 29GB of RAM at 0.50 tok/s" front-paged HN, one of at least three parallel threads this week on squeezing huge MoE models onto.

Enterprise DNA |
Extreme local inference of frontier-scale open models keeps escalating.

AI Pulse · Under the Radar

The play

Local inference appetite is high but speed is still too slow for production, revisit in six months when hardware catches up.

A thread showing how to run Kimi K3, a frontier-scale open model, on just 29GB of RAM at half a token per second hit the front page of Hacker News this week. It’s one of at least three similar threads in the last few days, all focused on cramming enormous mixture-of-experts models onto hardware you can buy at a store. The speed is barely usable, but the sheer volume of interest tells you something: a lot of people would rather own the inference, even if it’s slow, than rent it by the token.

This isn’t about performance. It’s about control and cost predictability. If you’re running a few hundred queries a month through an API, paying per token is fine. But if you’re prototyping something that might scale, or you want to experiment without watching a meter, or you just don’t want your prompts leaving your network, local inference starts to make sense. The fact that hobbyists are now squeezing models that would have needed a data centre two years ago onto a gaming rig changes the calculation for small operators.

What it means for you

You don’t need to run models on your own hardware today. But you should know the option exists and is getting cheaper fast. If you’re building something that depends on an API, ask what happens if usage spikes or if you want to move it in-house later. The gap between cloud and local is narrowing, and that changes how you think about vendor lock-in.

This is exactly the kind of trade-off we model when we build an AI command centre: when does it make sense to own the stack, and when do you rent? The answer isn’t always obvious, but it’s worth asking before you’re three months into a contract.

The thread is on Hacker News if you want to see the technical details.

Free daily email

Get this every morning.

This brief is one item from today's AI Pulse, the short daily read we run for ourselves on what is actually happening in AI. Subscribe free and it lands in your inbox each morning.

Free daily email

Subscribe to the daily AI Pulse

One short read every morning on what is actually happening in AI. Free.

One email a day. Unsubscribe any time.