Flinders University researchers have evaluated two next-generation reasoning large language models (LLMs)—o3-mini and DeepSeek-R1—and found that when asked to describe fictional patients with common medical conditions, the models frequently reproduced racial and gender stereotypes, indicating that advancements in AI reasoning do not inherently improve representational fairness.