Thomson Reuters spends $40M to build its own legal AI on Qwen

Thomson Reuters spends $40M to build its own legal AI on Qwen

Thomson Reuters has launched its first in-house AI language model, called Thomson, built on Alibaba's open Qwen. The company says it spent about $40 million on staff and computing power over more than two years to build it, a figure that dwarfs the more widely quoted $450,000, which covers only the final training run of the current version. Even the full $40 million leaves out the real capital behind the project: decades of content from Westlaw, Practical Law, Checkpoint and Reuters, plus the working hours of hundreds of domain experts.

The foundation is Alibaba's Qwen3.5-397B. Working with Imperial College, Thomson Reuters first retrained the Chinese base model for safety, ethics and political neutrality, producing an intermediate version named Snowdon after the mountain in Wales. From there the company ran pre-training on its own content, post-training with domain experts, and agentic reinforcement learning inside its own tool environments. So far less than 10 percent of the company's available content has gone into training. CTO Joel Hron says the team has changed the open source starting point close to half a dozen times already; research chief Jonathan Schwartz says the bigger result is less the individual model and more the model factory the company built to produce it.

The company's own benchmark numbers complicate its marketing claim that Thomson ranks among the best models in the world. On Stanford LegalBench, Thomson scores 0.823, trailing Gemini 3.1 Pro and GPT-5.5. On the Harvey Legal Agent Benchmark it sits just behind Opus 4.8. It leads on instruction following and on PrBench Legal, but falls off sharply on reasoning and especially coding. The comparison is also skewed by method: Thomson competes with test-time scaling turned on, while GPT-5.5 runs without a reasoning mode. On the company's own in-house Deep Research benchmark, with web access alone, Thomson scores 0.53 on factual accuracy against 0.65 for GPT-5.4. Only when given access to the company's own content does Thomson edge past GPT-5.4, 0.83 to 0.82. Evaluation lead Andrew Bean says that with web access alone, Thomson is within the scope of the other models but certainly not the leader yet, and that a big uplift comes from being able to train on and practice with the company's own tools, something outside providers cannot do. Notably, GPT-5.4 improves almost as sharply when given the same proprietary content, meaning data access is doing nearly as much work as the specialized training itself. The company did not test Thomson against newer models than the ones listed.

Thomson Reuters gives three reasons for building its own model rather than fine-tuning a frontier model from OpenAI or Anthropic. First, economics: Schwartz says standard fine-tuning techniques tend to have a strong tendency to degrade general capability, and fine-tuning still leaves a company locked into the provider for inference costs and roadmap; a smaller in-house model pays off on high-volume work like document review. Second, data: training inside the company's own tools, such as Westlaw, is where the performance gains actually come from, and the company will not hand that access to an outside provider. Third, a compounding effect that Hron compares to renting a house versus buying a house: every expert review during a product update becomes training data the company owns, while value built on a third-party model evaporates at the provider.

At launch, Thomson takes over the Tabular Analysis feature inside CoCounsel Legal, where a smaller, cheaper model makes economic sense; the product remains multi-model and administrators can switch between models. Thomson is not meant to act as an orchestrator, only to handle subtasks such as citation checking, and the company says customer data does not go into training. A smaller version is coming to Hugging Face as an open-weight model under a non-commercial license, with a technical report and developer portal to follow. Hron says early, still non-binding talks are underway with law firms about direct licensing.

Key facts

  • Thomson Reuters spent about $40 million on staff and computing power over more than two years to build Thomson, its in-house legal AI model built on Alibaba's Qwen3.5-397B; the often-quoted $450,000 figure covers only the final training run.
  • Less than 10 percent of the company's available proprietary content has gone into training so far.
  • On the company's own benchmarks, Thomson trails Gemini 3.1 Pro and GPT-5.5 on Stanford LegalBench (0.823) and sits just behind Opus 4.8 on the Harvey Legal Agent Benchmark, though it leads on instruction following and PrBench Legal.
  • Thomson only edges past GPT-5.4 on the company's Deep Research benchmark (0.83 vs 0.82) when given access to Thomson Reuters' own content; with web access alone it trails GPT-5.4, 0.53 to 0.65.
  • The model debuts in the Tabular Analysis feature of CoCounsel Legal for document review, with a smaller open-weight version coming to Hugging Face under a non-commercial license.

Why it matters

Thomson Reuters is a concrete data point in the build-versus-rent debate that every company with valuable proprietary data is having. Its case shows that a competitive, specialized model can be built for about $40 million by starting from an open-source foundation like Qwen and pairing it with exclusive domain data, rather than paying frontier labs indefinitely for inference and staying dependent on their roadmap.

Who it affects

Thomson Reuters customers using CoCounsel Legal and its document review workflows are the first to see Thomson in production. The result also matters to other companies weighing whether to fine-tune a frontier model from OpenAI or Anthropic versus training their own, and to open-source model providers like Alibaba, whose Qwen is now the base of a commercial legal product.

How to use it

Thomson enters CoCounsel Legal first through the Tabular Analysis feature, handling narrow, high-volume subtasks such as citation checking rather than acting as the product's orchestrator; the platform stays multi-model and administrators can switch which model runs a given feature. Customer data is not used in training. A smaller version of Thomson is headed to Hugging Face as an open-weight model under a non-commercial license, with a technical report and developer portal to follow, and Thomson Reuters says early, non-binding talks are underway with law firms about direct licensing.

How solid is it

The evidence is mixed and comes entirely from Thomson Reuters' own benchmarks. Thomson trails Gemini 3.1 Pro and GPT-5.5 on Stanford LegalBench and sits just behind Opus 4.8 on the Harvey Legal Agent Benchmark, while leading on instruction following and PrBench Legal and falling off sharply on reasoning and coding; the comparison is also skewed because Thomson runs with test-time scaling while GPT-5.5 does not use a reasoning mode. On the company's in-house Deep Research benchmark, Thomson trails GPT-5.4 on factual accuracy with web access alone (0.53 vs 0.65) and only edges ahead once given access to the company's own content (0.83 vs 0.82). The company has not tested Thomson against newer models than the ones it compared against.

Risks and caveats

The company's own numbers undercut its claim that Thomson ranks among the best models in the world: its lead over GPT-5.4 is thin and contingent, not decisive. GPT-5.4 improves almost as much as Thomson does when given the same proprietary content, which suggests that access to Thomson Reuters' own data, not the specialized training itself, is doing most of the work. With less than 10 percent of available content used so far, that balance could shift in either direction as training continues. Licensing talks with outside law firms remain early and non-binding.

“you are building equity in something that you own for the long-term, and that compounds over time”

— Joel Hron, CTO of Thomson Reuters