Back

Gaming the peer review system: a sophisticated review mill in medicine highlights the need to ensure reviewer integrity

Oviedo-Garcia, M. A.; Aquarius, R.; Bishop, D. V. M.

2025-10-23 obstetrics and gynecology
10.1101/2025.10.20.25338343 medRxiv
Show abstract

BackgroundA review mill is a network of researchers who game the peer review system to apparently boost their citations. Members write generic review reports containing suggestions for citations to the work of those in the review mill. Here we report compelling evidence for a review mill in the field of gynecologic oncology. MethodsThe first example of a peer-review report from a review mill was observed by chance. It contained boilerplate comments, such as: "Methodology is accurate and conclusions are supported by the data analysis" as well as suggestions that specific PubMed IDs be cited. We searched the internet using Google for review reports using the same boilerplate. We coded all review text to quantify similarities between review reports and compiled a list of citations suggested by reviewers. For 59 of 119 nonanonymous reviews in this target set, we identified a second review of the same article to act as a comparison. FindingsWe identified a set of 195 review mill reports that shared verbatim or highly similar boilerplate text from 170 targeted articles. 186 reports suggested citing at least one article which was co-authored by the reviewer or another member of the mill. Authors of 142 articles complied with some or all suggestions for citation. Nine of the reviewers in this review mill contributed 5 or more reviews; four had acted as editors for articles in the target set, and five had prolific peer review histories on Web of Science. Boilerplate text and self-citation recommendations were rare in the comparison reports. InterpretationReview mills threaten both the scientific record and patient safety when clinically relevant articles are improperly scrutinized during peer-review. We recommend that publishers adopt open peer review and transparently report the editors responsible for handling papers, to make it easier to detect review mills.

Published in Accountability in Research · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.