Invisible Text Injection: The Trojan Horse of AI-Assisted Medical Peer Review
Choi, B.; Jun, T. J.; Sung, J. W.; Park, I.; Lee, J.-M.; Cho, S. I.; Park, H. J.; Lee, R. W.; Suh, J.
Show abstract
Key PointsO_ST_ABSQuestionC_ST_ABSAre large language models robust against adversarial attacks in medical peer review? FindingsIn this factorial experimental study, invisible text injection attacks significantly increased review scores and raised manuscript acceptance rates from 0% to nearly 100%, while also significantly impairing the ability of large language models to detect scientific flaws. MeaningEnhanced safeguards and human oversight are essential prerequisite for using large language models in medical peer review. ImportanceLarge language models (LLMs) are increasingly considered for medical peer review. However, their vulnerability to adversarial attacks and ability to detect scientific flaws remain poorly understood. ObjectiveEvaluate LLMs ability to identify scientific flaws in peer review and their robustness against invisible text injection (ITI). Design, Setting, and ParticipantsThis factorial experimental study was conducted in May 2025 using a 3 LLMs x 3 prompt strategies x 4 manuscript variants x 2 with/without ITI design. We used three commercial LLMs (Anthropic, Google, OpenAI). The four manuscript variants either contained no flaws (control) or included scientific flaws in the methodology, results, or discussion section, respectively. Three prompt strategies were evaluated: neutral peer review, strict guidelines emphasizing objectivity, and explicit rejection. InterventionsITI involved inserting concealed instructions using white text on white background, directing LLMs to review with positive evaluations and "accept without revision" recommendations. Main Outcomes and MeasuresPrimary outcomes were review scores (1-5 scale) and acceptance rates under neutral prompts. Secondary outcomes were review scores, acceptance rates under strict and explicit reject prompts. We investigated flaw detection capability using liberal (detect any flaw) and stringent (detect all flaw) criteria. We calculated mean score differences by models and prompt types and used t-test and Fishers exact test for calculating P-value. ResultsITI caused significant score inflation under neutral prompts. Score differences for Anthropic, Google and OpenAI were 1.0 (P<.001), 2.5 (P<.001) and 1.7 (P<.001). Acceptance rates increased from 0% to 99.2%-100% across all providers (P<.001). Score differences were still statistically significant under strict prompting. Score differences were not significant under explicit rejection prompting, but flaw detection rate was still impaired. Using liberal detection criteria, results section flaw detection rate was significantly compromised with ITI, particularly in Google (88.9% to 47.8%, P<.001). Stringent criteria revealed methodology detection falling from 56.3% to 25.6% (P<.001) and overall detection dropping from 18.9% to 8.5% (P<.001). Conclusions and RelevanceITI can significantly alter the evaluation of medical studies by LLMs, and mitigation at the prompt level is insufficient. Enhanced safeguards and human oversight are essential prerequisites for the application of LLMs in medical publishing.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Transparency in peer review: Exploring the content and tone of reviewers' confidential comments to editors 95%
- Knowledge and motivations of training in peer review: an international cross-sectional survey 95%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 94%
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 92%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
- Developing an automatic system for classifying chatter about health services from Twitter: A case study for Medicaid 91%
Similar papers in this journal
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.