AI systems have to be pressure-tested in real-world clinical settings to ensure correct reasoning, accuracy and reliability.
Although it has always been well known that AI systems will get facts wrong occasionally in daily use, a new emerging problem is increasingly worrying clinical and technology leaders: AI convincingly appearing competent when it is completely incorrect and does not possess any real competence in an area at all. This phenomenon has been termed “cognitive spoofing” by industry experts, which refers to AI’s ability to confidently project expertise when it actually lacks the clinical context or judgement required to make a decision, especially in matters of care delivery. Because AI has the ability to create responses that are well polished and may seem backed by evidence, clinical leaders and physicians who put too much confidence into these systems can face mishaps in their daily workflows.
A study by Microsoft found that the latest frontier models were able to score exceedingly high on medical benchmarks and examinations. However, upon stress testing, the study found that these systems leveraged smart answering techniques rather than actual knowledge or sound reasoning: “Leading systems often guess correctly even when key inputs like images are removed, flip answers under trivial prompt changes, and fabricate convincing yet flawed reasoning. These aren’t glitches; they expose how today’s benchmarks reward test-taking tricks over medical understanding.” This means that although these systems score high on objective examinations, they may by no means be ready for actual clinical settings or stresses. As the study describes, clinical benchmarks for AI evaluation often emphasize correctness rather than reasoning or how that correct answer was reached. However, if this approach is pressure-tested in real time settings, it could lead to significant issues: “Medical readiness is a multidimensional construct. In real-world settings, models must tolerate missing or noisy data, justify their decisions in a manner clinicians can understand, and reason across time, modality, and context. Performance must be not only accurate but also reliable, interpretable, and safe under uncertainty.”










