HelpMum Africa, a health-tech organization dedicated to reducing maternal and infant mortality, today announced the release of HelpMum MamaBench, a counterfactual benchmark built specifically to evaluate how AI models reason about maternal and paediatric diagnosis not just whether they get individual answers right, but whether that reasoning actually holds up when a case changes in a clinically meaningful way.
The benchmark is fully open-source and available now, alongside a research paper detailing its findings on some of today’s most advanced AI models.
Why It Matters
Most existing medical AI benchmarks test one clinical question at a time, in isolation. That approach can hide a dangerous blind spot: a model may answer correctly on a standard case, then fail completely when a single decisive clinical detail shifts — even though, on paper, its accuracy score looks strong.
HelpMum MamaBench was built to expose exactly that gap. It consists of 434 clinical narratives, arranged in 217 matched pairs, spanning 371 maternal and paediatric conditions. Every narrative was individually authored from scratch by HelpMum Africa’s Medical team, written in first-person, patient-reported language rather than adapted from existing datasets. Each pair links a base clinical case to a counterfactual version, a case with the fewest possible changes needed to shift the correct diagnosis, testing whether an AI system’s reasoning is genuinely sound, or simply pattern-matching its way to a right answer.







