Thomson Reuters builds its own LLM, two years after buying Safe Sign

LLMsAI Agents
Thomson Reuters (TRI) CEO Steve Hasker On New CoCounsel, Its Own LLM, And The Widening Gap Between AI Adoption And Real Value

londoninsider.co.uk · News coverage photograph, editorial use approved

The Core · TL;DR

  • Thomson Reuters launched Thomson 1.0, an open-weights LLM trained on its own legal and business data, including Westlaw and Practical Law content
  • The model comes from the Safe Sign team, a Cambridge DeepMind spinout acquired in August 2024, after roughly 20 months of development and about $40 million in cost
  • Thomson Reuters says the model matches or beats frontier models on general and legal-specific benchmarks despite using less than 10% of its legal data
  • First deployment is inside Tabular Analysis in CoCounsel Legal, with a wider release to law firms and corporate legal teams planned

Thomson Reuters has released Thomson 1.0, an open-weights large language model trained on the company's own legal and business data rather than fine-tuned on top of a third-party foundation model. The launch marks the first commercial output of the team behind Safe Sign, a Cambridge University spinout founded by former Google DeepMind scientists that Thomson Reuters acquired in August 2024.

The project took roughly 20 months from acquisition to release and cost around $40 million, according to the company. Notably, Thomson 1.0 was trained on less than 10% of Thomson Reuters' total legal data holdings, alongside proprietary products such as Westlaw and Practical Law, suggesting the value came less from data volume than from the quality and structure of the legal corpus used.

Built for dense legal text, not just chat

Thomson Reuters says internal evaluations show Thomson 1.0 matching or beating frontier general-purpose models on both broad benchmarks and legal-specific tasks. The company points specifically to gains in instruction-following and in navigating dense, jargon-heavy legal documents, the kind of long, cross-referenced text that trips up generalist models trained mostly on web-scale data.

That framing puts Thomson 1.0 in a growing category of vertical models built by domain incumbents who already sit on proprietary datasets that outside labs cannot access. For a company whose core business is legal research and workflow software, a model tuned to statutes, case law, and contract language is a more defensible product moat than simply wrapping GPT-class APIs.

First stop: Tabular Analysis in CoCounsel

Thomson 1.0's initial deployment is inside Tabular Analysis, a feature within Thomson Reuters' CoCounsel Legal product. The company says a broader rollout to law firms and corporate legal departments is planned for an upcoming release, though no firm date has been given.

Building the model in-house also gives Thomson Reuters more control over cost, data governance, and update cycles than licensing a third-party LLM, an increasingly common calculation among enterprises with large proprietary datasets and regulatory sensitivities around client data.

Whether Thomson 1.0's benchmark edge holds up once it reaches a wider base of law firms and in-house legal teams will be the real test. For now, the model represents a two-year bet that owning the entire stack, from acquisition to training to deployment, pays off in a market where legal AI accuracy carries direct liability consequences.

WK

WAKIB Editorial Team

This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.

Subscribe to Newsletter

Get a weekly summary of the most promising AI research and tools delivered to your inbox.

Telegram Channel

Join our active community on Telegram for real-time tracking of AI models and trends.

Join us on Telegram