What a “frontier model” is, and how we built one for less
What the industry calls a “frontier model” is, in plain terms, a system at the leading edge of capability, the bar every other model is measured against. Until now, that bar has been set by a small number of heavily funded labs, each spending billions of dollars and years of infrastructure to reach it.
We took a different path. With Thomson, our first proprietary large language model, we built a system that competes with those frontier models, but did so with fewer than three dozen people, in three months from first experiments, with a final training run estimated at under $450,000 in GPU costs, specialized using Thomson Reuters’ authoritative professional content and expert-created data, while preserving broad capabilities.
That efficiency is the headline. What matters more is what’s behind it: how do we know any of it is true?
The answer is in the technical report we published alongside the launch of Thomson. It’s built to the standard we call Fiduciary-Grade AI™: AI for professionals with duties of care, where “almost right” is not good enough. The report is worth understanding at a summary level even if you never open the full PDF, because the methodology is what makes the claim credible.
“Thomson is competitive with the world’s leading frontier models despite being a fraction of their size and cost to train and operate.”
Chief Technology Officer, Thomson Reuters
Thomson: The technical report
Review our findings and details of the methods, data, and evaluations
Read the full report ↗
Continual learning
The most interesting idea in the report isn’t the benchmark scores. It’s how the model was built. Thomson wasn’t trained from scratch, which would have required the billions in compute that frontier labs spend. But it also wasn’t simply fine-tuned on top of an existing model, which typically buys narrow domain gains at the cost of broader capabilities. Instead, we started with open-weight models and applied what the report calls Continual Learning: a training approach that materially reshapes the model across the entire training stack while deliberately preserving the capabilities it already had.
The result is what the report describes as a T-shaped performance profile: pronounced gains in the professional domains we targeted, alongside preserved and frequently improved performance in domains we didn’t. That last part is the surprise. Most domain adaptation trades breadth for depth; Continual Learning, as we applied it, largely avoided that trade-off.
The report makes a broader case that this approach is repeatable: a blueprint for other institutions to build competitive models from open-weight starting points, at a fraction of the cost previously imagined.
Two kinds of evidence
The report doesn’t rest on a single number. It evaluates Thomson two different ways, and the distinction is the most useful thing to take from it: one tests the model in the lab, under identical conditions; the other tests it in the room, against the messy way professionals actually ask questions. Having both is what makes the claim worth taking seriously.
- The first is a standard benchmark comparison: Thomson-1.0-Large measured against today’s leading models, under identical conditions, across legal, tax, journalism, safety, and general-purpose tasks. Thomson lands exactly one percentage point behind the top-scoring model on an overall, unweighted average (79.5 vs. 78.5), and it leads every model tested on two specific measures: instruction following and a political-neutrality evaluation. Those two measures matter more than their category names suggest. Specializing a model on dense professional content usually costs something elsewhere. Narrow training tends to buy domain performance by spending general reliability. Instruction following and neutrality holding up, rather than slipping, is evidence the specialization was done with more care than brute force. The report breaks all of this down further in a full comparison table for anyone who wants the underlying numbers.
- The second kind of evidence is closer to how the model gets used day to day: a blind preference study comparing complete systems, not just models. Subject-matter experts rated thousands of real conversations without knowing which system had produced which answer. In legal conversations specifically, expert raters preferred Thomson’s answers to those of several named frontier systems a clear majority of the time. This is notable partly because this was a system comparison, not just a model comparison: Thomson had access to our legal tools and Reuters news, while the external models had broader web access. The report lays out the exact win rates against each named system, along with a separate, tighter set of results for general (non-legal) conversation, worth a look if you want the granular comparison. The per-system numbers is what a technical reader will want to scrutinize.

Where it loses, and why that’s the point
The most convincing part of the report is where it says, plainly, that Thomson isn’t ahead everywhere. For a skeptical technical audience, this matters more than any win rate: a launch document that volunteers a loss is one you can trust to report its wins honestly.
There’s one domain where Thomson-1.0-Large scores below its own starting model, attributed to mild forgetting during specialization and explicitly flagged as not a target area for this release. General-purpose reasoning, similarly, trails the strongest proprietary systems rather than leading them.
That’s worth sitting with. The version the data actually supports is narrower and more useful: Thomson is built and tuned for professional, domain-specific work, competitive with frontier systems on the tasks it was built for, and not represented as the best at everything. For the legal, tax, and compliance professionals this model is meant to serve, that’s the claim that should matter: not whether Thomson tops a general leaderboard, but whether it holds up on the specific kind of question they actually ask it.
The report’s constituent-benchmark tables are where that distinction gets precise, if you want to see exactly which tasks it covers and where the boundaries of the claim sit.
An audit, not an advertisement
It’s tempting to treat a technical report as supporting material for the announcement. It’s more accurate to treat it as the actual point.
The Fiduciary-Grade AI standard, the one this report exists to back up, only holds up if someone outside this building can check it, which is why the evaluation methodology carries as much weight here as any single score.
The two kinds of evidence only matter if they can be checked by someone other than us. The reason it holds up is that the evidence is open to outside scrutiny. The evaluations use publicly recognized benchmarks, re-implemented in a common harness so that any model is tested the same way. The report’s training record supports reproducibility and post-hoc analysis, so that for any released checkpoint the constituent datasets and training runs can be recovered exactly. And the smaller model is released as an open weight on Hugging Face for anyone to inspect directly.
Why SovereignAI matters
The report’s argument goes beyond building one strong model efficiently. Its larger stance is that organizations can own and control more of the AI stack than most assume: the model itself, proprietary data and tools, governance and values, infrastructure, and the economics of deployment. The report calls this SovereignAI, and frames it as a spectrum rather than a binary. Thomson doesn’t achieve full sovereignty on every axis, but it demonstrates meaningful progress across all of them, on a budget that makes the path viable for a much wider range of institutions than the current frontier-lab model implies.
If any of this is going to inform how you evaluate Thomson for your own work, read the report. The full benchmark tables, the preference-study breakdowns by system and by domain, the methodology behind each, and the open-weight model on Hugging Face are all there, worth reading firsthand rather than taking on faith.
Thomson
The purpose-built, proprietary LLM, engineered for high stakes professional work
Learn more ↗
