In a cinder-block clinic in one of Rwanda's rural districts, a community health worker unlocks her phone, opens a chat window, and types a question that, two years ago, she would have been forced to answer alone. A child has a fever that has not broken in three days. The nearest doctor is hours away by road, and the road, in April, is mostly mud. She describes the symptoms in Kinyarwanda, then in English, then in the awkward hybrid that her training has taught her the machine prefers. A few seconds later, the model replies. It is confident. It suggests a differential diagnosis, a likely cause, a set of next steps. The worker reads it twice. Then she makes a decision.
Multiply that scene by thousands. Multiply it again by the 101 community health workers who, in a study published in Nature Health on 6 February 2026, submitted 5,609 real clinical questions across four Rwandan districts to five different large language models. Multiply it by the 58 physicians in Pakistan who, in a parallel randomised controlled trial published in the same issue, were handed GPT-4o and twenty hours of training in how to argue with it, and whose diagnostic reasoning scores then jumped from 43 per cent using conventional resources to 71 per cent with the chatbot in the loop. By the researchers' own account, the large language models did not merely match the local clinicians. They beat them. Across every metric the team measured, the models won.






