Drastic changes in collaboration networks and publication patterns in research using the CDC WONDER dataset
Maupin, D.; Suchak, T.; Sengupta, A.; Marra, M.; Geifman, N.; Spick, M.
Show abstract
The growth of generative AI and easily available Open Access health datasets has transformed researcher productivity, leading to an explosion in publications that has in part been attributed to paper mills (organisations that provide manuscripts for payment) and other unethical actors. These entities are not, however, homogenous, and have a range of products and target markets. While the demand from China has received much attention, here we provide a case study of CDC WONDER, a dataset that has been exploited by a network of researchers reporting affiliations in Pakistan, the United States and the UK, potentially linked to medical residency driven demand from junior clinicians or trainees. The number of publications using CDC WONDER grew from 88 in 2021 to 1223 in 2025. Over the same time period, the proportion of papers reporting at least one author from Pakistan grew from 0.5% in 2021 to 27.2% in 2025, with unusually extensive collaboration networks. In some cases these works featured over 15 co-authors, often including representation from Western institutions, but in spite of this high level of resourcing only resulted in straightforward analyses of well-described conditions using publicly available data. The majority of these outputs additionally show evidence of being produced from a template, with formulaic titles and identical methods, for example using the same statistical model and platform (Joinpoint regression). Identifying papers produced by fast-churn workflows is essential to protect the integrity of the scientific literature from being flooded with low-quality research. This can be achieved through more proactive desk rejection of misleading and formulaic mass-produced submissions, and through better understanding of which use cases are appropriate for different Open Science resources. With the growing capabilities of AI to mass produce research, education will be essential to assist critical appraisal and preserve the benefits of Open Science.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Diversity and inclusion: A hidden additional benefit of Open Data 95%
- Artificial Intelligence's Contribution to Biomedical Literature Search: Revolutionizing or Complicating? 92%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 96%
- The Rise of Open Data Practices Among Bioscientists at the University of Edinburgh 96%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 94%
Similar papers in this journal
Similar papers in this journal
- Publishing at any cost: a cross-sectional study of the amount that medical researchers spend on open-access publishing each year 94%
- A systematic examination of preprint platforms for use in the medical and biomedical sciences setting 93%
- Comparison of preprints and final journal publications from COVID-19 Studies: Discrepancies in results reporting and spin in interpretation 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.