JudgeGPT lifts case resolutions 6.3% in Pakistan trial, quality holds up

JudgeGPT lifts case resolutions 6.3% in Pakistan trial, quality holds up

Pakistan's courts carry a backlog of 2.26 million cases and have fewer than two judges per 100,000 people, against 22 in the EU and eight in Brazil. To test whether AI could help, economist Sultan Mehmood of the New Economic School in Moscow and coauthor Elliott Ash of ETH Zurich, working with the judiciary, built JudgeGPT: a tool combining OpenAI's GPT-4 with retrieval-augmented generation over 128,292 Pakistani judicial opinions and 943 statutes (rounded to nearly 130,000 in the researchers' first mention), returning footnoted links to the cases and laws it draws on. Commercial AI chatbots, the team found, performed poorly on Pakistani legal queries and hallucinated case law frequently, which is why they built a tool tailored to the local context instead.

Starting in 2024, JudgeGPT was offered to 1,559 trial judges, roughly half the country's justices. Access alone was not the intervention: 1,197 judges went through six 90-minute Zoom training sessions run by Ash, covering how large language models work, their limits, bias and hallucination risks, and how to verify output. Another 180 judges received only general technology training with no JudgeGPT specifics, and a final group got no training at all.

By the time 487 judges had completed the full program, the median district saw resolved cases jump 6.3%, and the effect grew larger in districts with more trained judges. Appeal rates fell slightly rather than rising, which the researchers read as evidence that faster resolution was not producing sloppier decisions. Because having legal experts review large numbers of judgments was not feasible, the team instead had OpenAI's GPT-5-mini compare paired judgments by the same judge from before and after training; the model picked the post-training judgment 59% of the time. Two experienced Pakistani lawyers separately reviewed 90 of those judgment pairs and agreed with GPT-5-mini's choice 70.6% of the time, versus 73% agreement between the two lawyers themselves. The researchers do not report hallucination rates for JudgeGPT itself.

Training also determined whether judges kept using the tool. Judges who completed the full course logged in 56 times and sent 212 prompts on average over the study period, compared with 10 logins and 25 prompts for those given only generic training. Judges with no training tended to use the tool for about a month before dropping it entirely. Mehmood said the judiciary was more enthusiastic about the tool than the research team expected, given how badly delays were hurting litigants.

The researchers calculate that a trained judge resolved 38.5 more cases a month than the baseline, translating to roughly $38.50 saved in judicial costs for every dollar spent running the tool. Ash cautioned that these productivity and savings figures come from a nine-month period at the start of the trial, and that the underlying model and database have since been updated. One anonymous participating judge said their caseload has not dropped below 1,000 assigned cases in more than a decade, and that JudgeGPT now handles in seconds research and document-summarizing that previously took hours.

Outside commentators were cautiously positive but flagged open questions. MIT economist David Autor called large-scale field experiments in civil service rare and difficult to pull off, and described the 6.3% boost as credible, if not overwhelming, and likely to grow as the tool sees wider use. John Zeleznikow of La Trobe University said the study shows courts resolving cases faster but does not settle whether the underlying quality of justice has improved, since that goes beyond what appeal rates and pairwise judgment comparisons can show. The study also found that roughly a fifth of participants' prompts to JudgeGPT amounted to what the authors call substantial AI delegation: asking the tool what the best decision is, or to produce legal reasoning or write opinions with little input from the judge, though training reduced how often this happened. Ash said fixing hallucinations is less about smarter models than about attaching a model to a tool that can search and verify sources, but added that risks remain even with safeguards in place, and that judges should still be encouraged not to rely on the tool too much.

Key facts

  • JudgeGPT, built on GPT-4 with retrieval over 128,292 Pakistani judicial opinions and 943 statutes, was offered from 2024 to 1,559 trial judges, roughly half the country's justices.
  • By the time 487 judges had completed training, the median district saw resolved cases rise 6.3%, with appeal rates falling slightly rather than rising.
  • Trained judges (1,197, via six 90-minute sessions) logged in 56 times and sent 212 prompts on average, versus 10 logins and 25 prompts for the 180 judges given only general tech training.
  • Quality check: GPT-5-mini preferred post-training judgments 59% of the time; two lawyers agreed with the model 70.6% of the time versus 73% agreement with each other.
  • Researchers estimate roughly $38.50 saved in judicial costs per dollar spent running the tool, based on a nine-month early-trial window; about a fifth of prompts showed what the authors call substantial AI delegation.

Why it matters

Judges elsewhere have made headlines for using generative AI on the sly; this is described as the first major independent assessment of ongoing, sanctioned judicial AI use. Pakistan's courts faced a backlog of 2.26 million cases and fewer than two judges per 100,000 people, against 22 in the EU and eight in Brazil, so the pressure to find something that works was real. MIT economist David Autor called the study a rare large-scale field experiment in civil service and described its 6.3% productivity result as credible, though not overwhelming, and likely to grow with wider use.

Who it affects

Pakistani trial judges, 1,559 of whom were offered the tool starting in 2024, roughly half the country's justices, and the litigants waiting on a backlog of 2.26 million cases. The study was led by economist Sultan Mehmood of the New Economic School in Moscow with coauthor Elliott Ash of ETH Zurich, working in consultation with the judiciary. AI tools for judges are also already rolling out in Brazil and India, and U.S. law professor Eric Posner has separately compared LLM and human judgments in a single case study.

How to use it

JudgeGPT pairs GPT-4 with retrieval-augmented generation over 128,292 Pakistani judicial opinions and 943 statutes, returning footnoted links to the cases and laws behind its answers, aimed at legal research and drafting judgments. Usage tracked training closely: judges who took the full six-session, 90-minute-per-session Zoom course logged in 56 times and sent 212 prompts on average, against 10 logins and 25 prompts for judges given only generic technology training; untrained judges tended to use the tool for about a month and then stop.

How solid is it

The 6.3% jump in resolved cases is a median-district figure, measured once 487 judges had completed the program, and Ash noted the productivity and cost-saving numbers come from a nine-month window at the start of the trial, with the underlying model and database updated since. Quality was checked indirectly rather than by full expert review: GPT-5-mini preferred post-training judgments over pre-training ones 59% of the time, and two Pakistani lawyers who separately reviewed 90 of those pairs agreed with the model 70.6% of the time, close to the 73% agreement rate between the two lawyers themselves. The researchers do not report hallucination rates for JudgeGPT itself.

Risks and caveats

John Zeleznikow of La Trobe University said the study shows courts moving faster but does not resolve whether the underlying quality of justice improved, since that sits outside what appeal rates and pairwise judgment comparisons can capture. The authors found that roughly a fifth of participants' prompts amounted to what they call substantial AI delegation, asking the tool what the best decision is or to produce legal reasoning or draft opinions with little judge input, though training lowered how often this occurred. Ash said risks persist even with safeguards in place, and argued that alongside technological guardrails, judges still need to be discouraged from relying on the tool too heavily.

“We do find an increase in cases resolved, and we don’t find any corresponding decrease in decision quality,”

— Sultan Mehmood, economist, New Economic School in Moscow