Quantifying new threats to health and biomedical literature integrity from rapidly scaled publications and problematic research
Spick, M.; Onoja, A.; Harrison, C.; Stender, S.; Byrne, J.; Geifman, N.
Show abstract
Background and ObjectivesThe last three years have seen an explosion in published manuscripts analysing open-access health datasets, in many cases presenting misleading or biologically implausible findings. There is a growing evidence base to suggest that this is due in part to AI-assisted and formulaic workflows, and publishers are responding by discouraging submissions employing open-access health datasets. MethodsHere we employ a scientometric analysis to investigate which datasets have seen publication rates deviate from previous trends, especially where this coincides with changes to author geographical origins and increases in formulaic titles. ResultsAcross 36 datasets we identify nine showing hallmarks of paper mill exploitation (FAERS, NHANES, UK Biobank, FinnGen, the Global Burden of Disease Study, MIMIC, CHARLS, CDC WONDER, and TriNetX). These nine datasets had, in 2025, a combined publication count of 23,005 indexed in the OpenAlex database. This represents an excess of 11,577 publications above the AutoRegressive Integrated Moving Average (ARIMA) forecast trend, and is a 3.0x fold change on the 7,655 publication count for these nine datasets in 2022. We also identified a notable difference in the fold change for China (4.2x) versus the rest of the world (1.9x) and an increase in formulaic titles. ConclusionsThese findings highlight potential risks to research integrity in areas such as public health and drug safety, and especially to the accessibility and interoperability principles central to Open Science and FAIR data practices. We argue that permissive open-access data policies naturally facilitate exploitative workflows, and that these findings add to the case for the safeguarding mechanisms to preserve the goals of Open Science
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Publishing at any cost: a cross-sectional study of the amount that medical researchers spend on open-access publishing each year 94%
- A systematic examination of preprint platforms for use in the medical and biomedical sciences setting 94%
- Protocol for Development of a Reporting Guideline for Causal and Counterfactual Prediction Models 91%
Similar papers in this journal
- The Rise of Open Data Practices Among Bioscientists at the University of Edinburgh 95%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 94%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.