Jul 21, 2026
AI

JudgeGPT trial in Pakistan reports $38.50 return per dollar spent

A randomized trial across 118 Pakistani courts found AI access raised case clearances only when judges received targeted training.

Renata Fuchs

By Renata Fuchs · Policy Reporter

· 3 min read

JudgeGPT trial in Pakistan reports $38.50 return per dollar spent
Photo: The Decoder

Researchers from ETH Zurich, Imperial College London and the New Economic School tested an AI assistant called JudgeGPT across Pakistan’s trial courts and reported an estimated $38.50 in savings for every dollar spent. The result matters for government technology buyers because the gains came from a controlled rollout in a high-volume public institution, and because access to the model alone produced little benefit.

The randomized field experiment covered 1,559 judges in 118 courts, which the researchers described as roughly half of Pakistan’s trial court judiciary. Pakistan had 2.26 million pending cases at the end of 2024, according to the authors, with 82% sitting in trial courts. They also said the country has fewer than two judges per 100,000 residents, compared with 22 in the EU and 30 in England and Wales.

JudgeGPT was built on OpenAI’s GPT-4 for Pakistani trial courts. The system used retrieval augmented generation against a database of 129,235 documents, including 128,292 court rulings and 943 Pakistani laws. For each query, it selected 10 relevant passages and generated an answer with citations.

Training, not access, drove usage

The researchers divided judges into three groups. One group received JudgeGPT plus targeted training: six 90-minute sessions over three weeks, held after court hours and taught by ETH Professor Elliott Ash. The sessions covered appropriate use cases, limitations and verification of outputs. A second group received JudgeGPT and a general seminar on technology and law. A control group received the same general seminar without access to JudgeGPT.

The difference in usage was large. Judges who received targeted training used JudgeGPT four times as much as those who received only the general seminar, according to the researchers. After 40 weeks, trained judges averaged nearly 60 logins and more than 200 prompts. Judges with access but no targeted training averaged about 20 logins and fewer than 50 prompts.

Districts with more trained judges cleared more cases. At moderate exposure, the researchers estimated about 1,848 additional resolved cases per district per year, equal to a 6.3% increase. Even districts in the bottom quartile of exposure cleared roughly 616 extra cases.

Quality measures did not show a trade-off

The study did not find evidence that the productivity gain came at the expense of judgment quality. Appeal rates per 1,000 resolved cases fell slightly, according to the researchers, while judges reported no change in hours worked or work-life balance. The $38.50 return estimate was based on the cost of hiring enough additional judges to produce the same increase in output. The researchers said conservative assumptions still implied at least $10 in savings per dollar invested.

A review of about 4,000 judgments found more AI-flagged text, which the researchers expected. Readability, length and the number of legal arguments were stable. An LLM-based quality assessment, validated by two Pakistani lawyers, showed a modest improvement: rulings from trained judges were preferred in 59% of pairwise comparisons, compared with 42% for the control group. The study found no evidence that JudgeGPT increased gender or religious bias in judicial language.

Use shifted toward lower-risk tasks

An analysis of anonymized chat logs from about 1,500 judges found that legal research, editing and text generation were the most common uses. About 60% of queries asked for information about laws, procedures or legal concepts.

Training changed behavior. Trained judges used JudgeGPT more often for editing and summarization, tasks where language models are generally more dependable, and less often for broad legal questions where hallucination risk is higher. About one-fifth of requests involved what the researchers called substantive AI delegation, such as asking the system to evaluate decisions or draft reasoning on its own.

The researchers said the findings support AI as a productivity aid for courts, rather than a replacement for judges. The less comfortable conclusion for vendors is also clear: procurement without workflow-specific training may leave most of the value unrealized.

This story draws on original reporting from The Decoder.

More from AI

All AI →