Flinders University researchers have evaluated two next-generation reasoning large language models (LLMs)—o3-mini and DeepSeek-R1—and found that when asked to describe fictional patients with common medical conditions, the models frequently reproduced racial and gender stereotypes, indicating that advancements in AI reasoning do not inherently improve representational fairness.

"Large language models have the potential to transform health care but risk exacerbating health disparities if they perpetuate biases," says lead researcher Joshua Docking, from Flinders University's College of Medicine and Public Health.

Researchers have previously demonstrated potential racial and gender biases in clinical vignettes generated by GPT-4, including overrepresentation of Black patients in stereotypical medical conditions. Since then, next-generation reasoning LLMs have emerged, offering improved reasoning capabilities and demonstrating superior benchmark performance.

"Whether these advances reduce representational bias in health care remains unknown, so this study evaluated whether reasoning LLMs exhibit racial and gender biases in generated clinical content."

Newer models showed no gains