The test ran 200 cases across 16 models and compared them against a panel of radiologists. Human experts scored 988.7 out of a possible 2,000 points. The best AI model hit 758.
Human radiologists outperformed every tested AI model on the primary metric, which combines accuracy with the confidence level of each answer. | Image: Crash Lab
Honest silence beats overconfident guesswork
The scoring system rewards honesty and punishes overconfidence. Get it right with high confidence, and you earn full points. Get it wrong while claiming high confidence, and you lose a matching number. Answer "I don't know," and you score zero but don't lose anything. A model that guesses confidently drops in the rankings even if its raw hit rate looks decent.
The study tackles a point recently raised by this highly cited paper: as long as benchmarks only reward accuracy, AI models are trained to guess. In medicine, a confident misdiagnosis is far more dangerous than an honest admission of uncertainty.









