Researchers have built a tool to uncover hidden issues in the massive datasets used to train medical AI. The tool searches training data for subtle patterns that could lead AI models to incorrect conclusions, potentially jeopardizing patient care. The work aims to help researchers and regulators make artificial intelligence more reliable and trustworthy for real-world clinical use.

"The models that drive precision medicine learn to infer clinical outcomes from the data they're trained on. In many cases that works great, but it can also lead to interesting failures that aren't immediately apparent," said senior author Mathias Unberath, an expert in AI-assisted medicine at Johns Hopkins University. "The tool we developed gives you a clear understanding of which elements of metadata pose the greatest risks for your model to pick up bias."

The work, a collaboration between Johns Hopkins and the U.S. Food and Drug Administration, is published in npj Digital Medicine.

When models learn the wrong lesson

The "Clever Hans phenomenon" refers to a horse named Hans that became famous in the early 1900s for being able to do math, tapping out correct answers with a hoof. It turned out Hans couldn't do math but was quite good at reading his human's body language, which inadvertently revealed how he should respond.