Fine-Grained Detection of AI-Generated Writing in the Biomedical Literature
She, R.
Show abstract
Generative AI systems are rapidly being integrated into scientific workflows, yet the specific ways in which AI-generated prose appears in published literature remain poorly characterized. Here, we use Pangram, a transformer-based detector optimized for adversarial paraphrasing, to analyze full-length biomedical research articles from 13 major journals. Papers published in 2021-2024 showed almost no detectable AI-generated text, whereas manuscripts published in 2025 exhibited a sharp increase, with 12.4% containing at least one localized passage classified as AI-written. AI usage was highly nonuniform across authors and geography: 32% of papers originating from South Korean institutions and 26% papers from Chinese institutions contained AI-generated passages, compared to 7.4% from U.S. institutions. In a focused case analysis, six labs that published fully AI-generated manuscripts also produced additional papers with extensive AI-generated segments. Journals likewise differed, with high-selectivity venues rarely containing AI-authored prose, while high-volume journals accounted for most AI-positive manuscripts. Together, these findings provide the first detailed empirical map of how and where AI-generated writing is entering the scientific literature, underscoring the need for clear norms and policies governing the use of generative AI in scientific communication.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 93%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 90%
- Large language models improve transferability of electronic health record-based predictions across countries and coding systems 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.