Jul 19, 2026
AI

RadLE 2.0 finds radiology AI still guesses too confidently

Ashoka University’s CRASH Lab tested 16 AI models on 200 radiology cases and found humans still led when accuracy and confidence were scored together.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 3 min read

RadLE 2.0 finds radiology AI still guesses too confidently
Photo: The Decoder

Ashoka University’s CRASH Lab has released RadLE 2.0, a radiology benchmark that scores AI models not just on whether they identify a case correctly, but on whether they know when to defer to a doctor. Across 200 cases and 16 models, a panel of radiologists scored 988.7 out of 2,000 on the benchmark’s main confidence-weighted metric, while the top AI system scored 758.

The result is a useful check on a common AI sales claim in medicine: high raw accuracy does not mean a system is safe to run without supervision. RadLE 2.0, short for “Radiology’s Last Exam,” asks models to rate confidence from 0 to 4 and allows them to answer “I don’t know.” That design makes the benchmark less forgiving of systems that produce plausible answers while overstating certainty.

The benchmark penalizes confident errors

CRASH Lab’s scoring approach rewards correct, confident answers and subtracts points when a model is confidently wrong. A refusal to answer earns no points, but it also avoids a penalty. That distinction matters in clinical workflows, where a wrong answer delivered with certainty can create more risk than a handoff to a specialist.

The researchers connect the test design to a broader concern in AI evaluation: benchmarks that measure only accuracy encourage models to guess. In RadLE 2.0, that behavior can lower a model’s rank even if its hit rate looks competitive.

No model dominated every measure. According to the research team, Anthropic’s Claude Fable 5 led the primary reliability and safety metric. Google’s Gemini 3 Pro had the best raw accuracy. Meta’s Muse Spark 1.1 performed best on handover readiness, meaning it was strongest at recognizing when a case should go to a human radiologist.

The researchers said several systems would have scored higher if they had declined more cases. That weakness was especially visible among open-weight models and models built for medical vision-language tasks, which attempted nearly every case and were often wrong with medium or high confidence.

Accuracy is improving faster than calibration

The first RadLE release, published in September 2025, showed a wider gap: radiologists reached 83% accuracy, while the best model managed about 30%. Within three months, Gemini 3 Pro had passed the level of resident radiologists on accuracy, according to CRASH Lab. The new results suggest that raw diagnostic performance is moving quickly, while uncertainty estimation remains a harder problem.

That gap is commercially important. Patients are already uploading scans such as X-rays and MRIs to general-purpose chatbots, while vendors and investors continue to frame medical AI as close to autonomous diagnosis. CRASH Lab argues those claims are often based on anecdotes or simulations rather than clinical readiness. A separate April study of 21 then state-of-the-art models found they were not ready for unsupervised clinical use.

Other recent work has pointed in a more optimistic direction. The MIRA system for electronic health records and AMIE, a medical consultation system, were reported to keep pace with general practitioners in simulated settings. RadLE 2.0 does not contradict those findings directly, but it puts a narrower and tougher question to radiology models: can they tell when they should stop answering?

Radiology remains a hard automation target

Radiology has been a recurring target for AI replacement predictions since at least 2016, when Geoffrey Hinton argued that training radiologists should stop because deep learning would soon take over the job. That forecast did not hold. Radiologists remain in demand, and Hinton later walked back the prediction.

CRASH Lab plans to keep adding models to RadLE 2.0 and said a full scientific publication will include cost analysis and an error taxonomy. Those additions will matter for buyers comparing model performance, since inference cost and failure modes are often more useful than leaderboard placement alone.

The near-term readout for hospitals, founders and AI buyers is straightforward: radiology models are getting better at identifying cases, but many still present wrong findings with too much confidence. Until that changes, these systems look more like supervised tools than autonomous clinicians.

This story draws on original reporting from The Decoder.

More from AI

All AI →