Some of the artificial intelligence tools being used to predict a person’s risk of stroke or diabetes may rest on a shaky foundation: datasets whose true origins cannot be verified, according to new research from Queensland University of Technology (QUT) and the Australian Centre for Health Services Innovation (AusHSI).

The study, published in BMC Medicine, examined two widely downloaded health datasets hosted on Kaggle, an online platform for sharing datasets and machine-learning resources that markets itself as “the world’s AI proving ground.” Despite their popularity — the datasets have been used in 125 peer-reviewed studies — the researchers found they provide almost no information about where the data actually came from, how it was collected, or whether it represents real patients at all.

An Enormous Surprise

Lead author Alexander Gibson, from the QUT School of Public Health and Social Work and AusHSI, said the team was taken aback by what they found buried in datasets that have quietly shaped clinical AI research for years.

“It was an enormous surprise to come across something like this. These datasets exhibit unusual patterns that raise serious questions about their authenticity and suitability for clinical research.”

Alexander Gibson, Queensland University of Technology

The stakes are not merely academic. Three prediction models built on the data have evidence of use in clinical practice, one model was cited in a medical device patent, and the models collectively appear in 86 review articles — meaning flawed data may have quietly worked its way into decisions that affect real patients.

Using the internationally recognized TRIPOD+AI reporting framework, the researchers assessed the datasets against nine essential data-provenance criteria. The datasets scored zero out of nine.

“Prediction models built on data of unknown provenance have no place in clinical decision-making. Without trustworthy data, the outputs are unreliable and risk misleading clinicians and harming patients.”

Alexander Gibson

A Broader Warning About Fast-Churn AI Research

The authors are calling on journals, funders, and data repositories to strengthen requirements for disclosing where data comes from, and they recommend that the two Kaggle datasets be removed entirely to prevent further misuse. Seven articles that relied on the datasets have already been retracted for being unreliable, and the findings have prompted an update to the Collection of Open Science Integrity Guides.

“We’re seeing fast-churn research built on datasets that look scientific but lack the most basic transparency. Without stronger safeguards, unreliable models will continue to make their way into the literature, and potentially into practice.”

Alexander Gibson

The study also included QUT researchers Professor Adrian Barnett and Associate Professor Nicole White.


The study, “Evidence of unreliable data and poor data provenance in clinical prediction model research and clinical practice,” was published July 9, 2026, in BMC Medicine (DOI: 10.1186/s12916-026-04981-y).

Leave a Reply

Trending

Discover more from Scientific Inquirer

Subscribe now to keep reading and get access to the full archive.

Continue reading