Revealing the Paper Mill Iceberg: AI-Based Screening of Cancer Research Publications.
Scancar, B.; Byrne, J. A.; Causeur, D.; Barnett, A. G.
Show abstract
ObjectivesTo train and validate a machine learning model to distinguish paper mill publications from genuine cancer research articles, and to screen the cancer research literature to assess the prevalence of papers that have textual similarities to paper mill papers. DesignMethodological study applying a BERT-based text classification model to article titles and abstracts. SettingRetracted paper mill publications listed in the Retraction Watch database were used for model training. The cancer research corpus was screened by the model, using the PubMed database restricted to original cancer research articles published between 1999 and 2024. ParticipantsThe model was trained on 2,202 retracted paper mill papers and validated on independent data collected by image integrity experts. A total of 2.6 million cancer research papers were screened. Main outcome measuresClassification performance of the model. Prevalence of papers flagged as similar to retracted paper mill publications with 95% confidence intervals and their distribution over time, by country, publisher, cancer type, research area, and within high-impact journals (Decile 1). ResultsThe model achieved an accuracy of 0.91. When applied to the cancer research literature, it flagged 9.87% (95% CI 9.83 to 9.90) of papers and revealed a large increase in flagged papers from 1999 to 2024, both across the entire corpus and in the top 10% of journals by impact factor. Over 170,000 papers affiliated with Chinese institutions were flagged, accounting for 35% of Chinese cancer research articles. Most publishers had published substantial numbers of flagged papers. Flagged papers were overrepresented in fundamental research and in gastric, bone, and liver cancer. ConclusionsPaper mills are a large and growing problem in the cancer literature and are not restricted to low impact journals. Collective awareness and action will be crucial to address the problem of paper mill publications.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Full Publication of Preprint Articles in Prevention Research: An Analysis of Publication Proportions and Results Consistency 92%
- CONSORT-TM: Text classification models for assessing the completeness of randomized controlled trial publications 91%
- A natural language processing system for the efficient extraction of cell markers 91%
Similar papers in this journal
Similar papers in this journal
- Novel significant stage-specific differentially expressed genes in liver hepatocellular carcinoma 89%
- Deep learning-based tumor microenvironment segmentation is predictive of tumor mutations and patient survival in non-small-cell lung cancer 89%
- Gene networks and expression quantitative trait loci associated with platinum-based chemotherapy response in high-grade serous ovarian cancer 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.