As of today, July 27, 2026, Moonshot AI has published the full model weights for Kimi K3 on HuggingFace under a Modified MIT license. That makes Kimi K3 the largest open-weight AI model ever released: 2.8 trillion parameters, free to download, with the right to inspect, fine-tune, and self-host.
For business leaders, the question is no longer “is this coming?” It is here. And it changes things in ways that are both more nuanced and more immediately useful than most of the coverage suggests.
What Actually Dropped Today
Kimi K3 is a mixture-of-experts model. Despite the headline 2.8 trillion parameter count, it activates only around 50 billion parameters per token during inference, using 16 of its 896 expert layers at any one time. That architecture is what lets Moonshot claim real-world serving costs that compete with models a fraction its nominal size.
The weights ship in MXFP4 quantization. The download is roughly 594 gigabytes. Loaded into memory for serving, you need approximately 1.4 terabytes of fast GPU memory, and Moonshot recommends a minimum of 64 accelerators configured as a single pool. The model carries a 1-million-token context window and topped the Frontend Code Arena leaderboard on its initial release July 16.
The Part Most Coverage Is Skipping
Since K3’s launch, independent benchmarkers have run it through the AA-Omniscience hallucination benchmark. The results are worth knowing before you build anything on this model.
Kimi K3 hallucinated on approximately 51% of AA-Omniscience test cases. Its predecessor, Kimi K2.6, scored 39% on the same benchmark. That is a meaningful regression, in the direction you do not want, on a capability that matters enormously for production business use cases.
To be clear: Kimi K3 is genuinely impressive at coding tasks. It can generate, debug, and reason about complex code at a level that a few months ago would have required a closed frontier model. But code generation and factual retrieval are different cognitive tasks, and K3’s hallucination profile suggests it should be treated differently from models like Claude Fable 5 or GPT-5.6 for knowledge-retrieval and question-answering workloads.
Use it where accuracy can be verified downstream. Be cautious where it cannot.
The Data Sovereignty Question Is Now Real
When Kimi K3 launched on July 16, using it meant sending your data through Moonshot’s API, which runs in China. That creates two categories of risk. The first is data residency: business data leaving your jurisdiction. The second is legal: Chinese law creates data obligations for companies operating under it, and no vendor agreement eliminates that exposure.
Today changes the calculation. With the weights available, organisations with sufficient infrastructure can self-host K3 on their own servers, in their own jurisdiction, under their own data policies. That removes most of the API risk. What remains is a reputational and governance question: there was an April 2026 cross-user data breach from Moonshot’s infrastructure that the company never publicly addressed. How you weigh that as part of your AI vendor risk assessment is a judgment call. But the option to avoid the API entirely now exists.
What This Means for Business
You probably cannot run this directly. 64 accelerators and 1.4 terabytes of GPU memory is not infrastructure that most organisations have sitting around. The teams that can self-host immediately are inference providers, cloud platforms, and very large enterprises with dedicated AI infrastructure. For most businesses, you will access K3 through inference providers who run the weights on your behalf, at prices that will likely be significantly cheaper than the Moonshot API.
Your negotiating position with AI vendors just changed. A credible, free, frontier-class model changes the pricing dynamics for everyone in this market, whether or not you switch. Even if you stay on Anthropic or OpenAI, the existence of K3 as a viable alternative gives you leverage in conversations about pricing, terms, and data handling. Use it.
The coding use case is real, right now. If your development team is not already evaluating K3 through inference providers for code generation and review tasks, they should be. The benchmark lead on coding is genuine, and the economics of open-weight models mean that route will get cheaper as more providers compete.
Build with the hallucination rate in mind. Do not build customer-facing knowledge systems on K3 without a verification layer. The 51% hallucination rate on factual benchmarks is not disqualifying for all use cases, but it is disqualifying for some. Match the model to the task, and design your workflows accordingly.
The Bigger Picture
The Kimi K3 open weights release is part of a broader compression happening in the enterprise AI market. The gap between frontier closed models and the best open-weight alternatives is narrowing faster than most business leaders expected. That compression is good for buyers. It means the premium for proprietary AI keeps shrinking, and building on open infrastructure becomes more viable with every passing quarter.
At Enterprise DNA, we help organisations navigate exactly these decisions: which models to build on, how to structure data infrastructure, and where the real leverage in AI adoption lives right now.
If your team is working through an AI model strategy and wants a clearer view of the options, we would be glad to talk. Book a session with our Omni Advisory team to get oriented quickly.
Source
TECHi