Enterprise DNA

Omni by Enterprise DNA

Enterprise DNA Resources

Latest AI and industry news. Practical AI operating-system thinking for owners, operators, and teams doing real work.

220k+

Data professionals

Omni

AI agents and apps

Audit

Map the manual work

News Trending AI News

Thomson Reuters Built Its Own LLM for Just $40M

Thomson Reuters built a proprietary legal AI model on $40M and its own data. Here's what it signals for enterprise AI strategy.

Enterprise DNA | | via Thomson Reuters
Thomson Reuters Built Its Own LLM for Just $40M

Thomson Reuters just did something that most enterprises don’t think is possible: they built their own frontier AI model for $40 million.

Not $400 million. Not billions. Forty million dollars — and they trained it on data they already owned.

The model, called Thomson 1.0, launched on August 24, 2026. It’s the company’s first proprietary large language model, built from an open-source foundation and fine-tuned on decades of Westlaw, Practical Law, Checkpoint, and Reuters content. Thomson 1.0 will initially power the tabular analysis feature in CoCounsel Legal, the company’s AI assistant for legal professionals, with more features to follow.

The announcement is quietly significant — not just for legal tech, but for any business sitting on valuable proprietary data.

What Thomson Reuters Actually Built

Thomson 1.0 is not a general-purpose model. Thomson Reuters didn’t set out to compete with OpenAI or Google. They built a specialist — one trained on less than 10% of their total legal data — that performs better than general models on the specific tasks their customers need.

That’s the key insight buried in the press release: they didn’t need to train on everything. They trained on the right things.

The model was developed with the team from Safe Sign Technologies, a Cambridge AI startup Thomson Reuters quietly acquired in 2024. It’s built on an open-source foundation model, which kept costs manageable and gave the team a solid baseline to fine-tune rather than train from scratch.

The result is a model Thomson Reuters fully owns and controls — one that never sends sensitive client data to a third-party API.

Most business leaders assume building a proprietary AI model is something only trillion-dollar companies do. Thomson Reuters just showed that’s not true.

The $40M price tag will continue to fall. Open-source foundation models are getting better every quarter. Fine-tuning infrastructure is increasingly accessible. And the value of domain-specific models — models that actually understand your industry’s language, processes, and data — is only going up.

What Thomson Reuters had that most companies don’t is a clear data advantage. Decades of Westlaw case law. Practical Law guidance. Tax content from Checkpoint. Reuters news. That’s a corpus that no general model can replicate.

The question for every business leader is: what’s your equivalent?

If you run an accounting firm, you have years of client financials, reconciliation notes, and tax filings. If you run a logistics company, you have route data, freight rates, and supplier history. If you run a healthcare business, you have patient interaction patterns and clinical outcomes. All of that is training data — and most businesses treat it like a byproduct rather than an asset.

The CoCounsel Connection

The timing is interesting. Thomson Reuters also announced the next generation of CoCounsel Legal on August 20, 2026 — a fully agentic AI assistant built on Anthropic’s Claude Agent SDK that can reason, plan, and execute legal work at the level of a senior associate.

So the company is running a dual strategy: use best-in-class foundation models (Claude) for the orchestration layer, and use their own proprietary model (Thomson) for the domain-specific tasks where their data advantage is most valuable. This is not an either-or decision — it’s a layered architecture.

That’s a sophisticated approach that most enterprises haven’t figured out yet. You don’t have to choose between OpenAI and building your own model. You can use both, for different jobs.

What This Means for Business

For business owners and data leaders, the Thomson Reuters story is a template worth studying:

Proprietary data is now a production asset. If you’ve been collecting structured data about your operations for years, that data has AI training value. The cost of leveraging it is dropping fast.

You don’t need to train from scratch. Open-source foundation models give you 80% of the way there. Domain-specific fine-tuning on your own data gets you the rest — for a fraction of what frontier model training costs.

Reducing AI dependency is a legitimate strategy. Thomson Reuters built Thomson partly to reduce reliance on big tech AI providers. Control over your model means control over your costs, your data, and your roadmap.

Specialist models beat generalists on specialist tasks. A model trained on legal data will outperform GPT-4o on legal document review. A model trained on your ERP data will outperform it on your operational forecasting. The gap between specialist and generalist AI is where the next wave of enterprise value gets built.

Enterprise DNA works with businesses that are sitting on exactly this kind of underutilised data. The organisations winning with AI right now aren’t always the ones with the biggest budgets — they’re the ones who understand what their data is worth and build intentionally around it.

Thomson Reuters spent $40M to learn something every data-driven business should internalise: your most valuable AI asset might already be in your systems. You just haven’t trained on it yet.