Skip to main content
Rigor over momentum: Why courts must slow down to get AI right

Rigor over momentum: Why courts must slow down to get AI right

This is some text inside of a div block.
By:
Rabihah Butler,
Rabihah Butler
September 18, 2026
8 min
September 18, 2026
This is some text inside of a div block.
This is some text inside of a div block.

Courts face a painful paradox: The more pressure to adopt AI quickly that they face, the more they need to slow down and think strategically. A new webinar discusses how rigorous evaluation matters more than cutting-edge tools and shows court leaders what separates successful implementations from costly failure

Listen to this article

Key insights:

  • User satisfaction is not a measure of accuracy — Rigorous evaluation requires examining error rates, sources, and reproducibility.
  • Phased rollout reduces risk and builds confidence — Starting with expert users who can catch errors, then gradually expanding access as the tool proves itself, is more effective than trying to achieve perfection before launch.
  • Evaluation never stops — Success requires continuous assessment after deployment, not a one-time checklist.

Everyone wants to be part of the AI revolution and courts are no different. Judges and court staff see other jurisdictions adopting new tools, vendors arriving with compelling demos, and grant deadlines creating pressure to act quickly.

However, this pressure comes at a critical moment. More than two-thirds of judges and court professionals report staffing shortages, with caseloads increasing and staff working excessive hours. The fear of falling behind is understandable, yet this is exactly the point at which court professionals need to step back and think carefully about their AI implementation.

A recent webinar, Responsible by Design: Evaluating & Testing AI Tools for Accuracy and Reliability, underscored a central lesson that successful AI implementation depends less on the tool itself than on how rigorously courts evaluate and deploy it. The webinar is part of a series from the AI Policy Consortium for Law & Courts, a joint effort by the National Center for State Courts (NCSC) and the Thomson Reuters Institute (TRI).

Before you deploy: Five essential steps

As the webinar discusses, successful implementation requires understanding what you're solving and why an AI tool may be right for the job:

  1. First, understand the problem — You can't fix what you don't understand.
  2. Understand how that problem can be solved — Not every problem requires an AI solution. Sometimes process redesign, better forms, or staff training is the answer.
  3. Understand the available solutions — Look beyond features to understand each tool's fundamental architecture and real-world trade-offs.
  4. Understand the risks — AI tools carry specific risks, and they must each be examined and weighted.
  5. Understand the benefits — Be clear about expected gains and honest about whether they're worth the risks.

Once you've done that foundational work, vet any solution carefully. Many courts confuse user satisfaction with tool effectiveness.

During the webinar, Keith Porcaro, Assistant Clinical Professor of Law at Duke University School of Law, notes this critical distinction. "They might give a thumbs down to an answer that is correct but is not what they wanted to hear," Porcaro says. "They might give a thumbs up to something that looks really persuasive but is completely wrong."

Judges and court professionals should remember their role as arbiter of what happens in their court.

A thumbs-up-or-down survey tells you what users think about a tool, not whether it's actually accurate or consistent. Proper evaluation instead asks a plethora of different questions to evaluate errors, failures, and success.

Evaluation happens in three phases

The panelists frame evaluation as three distinct phases that court leaders own. First is the green light. Before spending money, they need to decide what they are actually aiming for, define their quality standards and harm floors (meaning, the outcomes you will never accept), and stress-test any demo with their own sample cases, running it multiple times to check that results are consistent.

Second, is the hill climb. This is the often-lengthy work of refining and retesting a promising tool until a scoped pilot is genuinely ready to launch.

And third, is of course, the ongoing maintenance. This requires a continued sampling and testing of real user sessions after launch, with a governance plan settled in advance that says who can pull the tool offline, how quickly, who pays for re-testing, and how the public is notified if something goes wrong.

Here's the hard truth, however: You can't delegate accountability. Someone within the court's organization needs to understand what good evaluation looks like and whether they've achieved it.

"Your job is to make sure that you're putting any proposed new project or almost a pilot development through sufficient review," explains Angela Tripp, former Program Officer for Technology at the Legal Services Corporation. "You don't need to get into the weeds yourself running tests, but you need to know what to ask for and how to interpret the answers."

Start small, prove It, scale It

Once a court has evaluated a tool they like, the natural impulse is to go big. This is when they need the most discipline.

Margaret Hagan, Executive Director of the Legal Design Lab at Stanford Law School, describes a more thoughtful approach. "Our first pilot of that kind is very internally scoped," Hagan says. "And then we might expand out to gradually let more junior people use it — but we're not aiming for general public users in our first pilot."

You can't delegate accountability. Someone within the court's organization needs to understand what good evaluation looks like and whether they've achieved it.

Indeed, courts can have real power in this process. They're not choosing between adopting a vendor's vision or staying frozen in fear. Judges and court professionals should always demand clarity on how tools were tested and what they actually do and ask tough questions about error rates, consistency, and reliability.

This information will allow them to better control the scope of their pilot to make evaluation realistic. It's also critical to build internal expertise to reduce dependence on vendors and to stay skeptical of impressive demos and other jurisdictions' success stories.

Judges and court professionals should remember their role as arbiter of what happens in their court.

Those courts that successfully integrate AI won't be the ones that move fastest. Rather, as the webinar shows, they'll be the ones moving deliberately, knowing what problems they're addressing, and starting small with pilots.

They'll also remember that the tool isn't the problem — the problem is rushing past the hard work of thinking clearly about what is needed and why, and whether this particular tool actually delivers.

Courts can genuinely improve access to justice through smart use of AI, however, that requires implementing the right tool, informed by rigorous evaluation, deployed deliberately, and managed continuously. As the webinar discusses, it is that discipline that separates courts that benefit from AI from those that stumble.

You can find out more about the work that NCSC is doing to improve courts here

Follow us on social

Have questions?

Get in touch with one of our solutions experts....
Thomson Reuters Institute logo
Rigor over momentum: Why courts must slow down to get AI right
Courts face a painful paradox: The more pressure to adopt AI quickly that they face, the more they need to slow down and think strategically. A new webinar discusses how rigorous evaluation matters more than cutting-edge tools and shows court leaders what separates successful implementations from costly failure
September 18, 2026
8 min
AI Governance
Rabihah Butler
Manager for Enterprise content for Risk, Fraud & Government
Thomson Reuters Institute
Headshot of Rabihah Butler
AI literacy
Access To Justice
Government Agencies
Human Side of AI
Webinars
NCSC
AI in courts
The tool isn't the problem: Courts should judge on merit, not technology
Courts grapple with AI revolution amid staffing crisis
Chatbots for justice: The impact of AI-driven tech tools for pro se litigants