Fully Automated Systematic Review Generation via Large Language Models: Quality Assessment and Implications for Scientific Publishing
McLaughlin, L.; Walz, M. S.; Arries, C.
Show abstract
Large language models (LLMs) are increasingly transforming scientific workflows, yet their application to rigorous evidence synthesis remains underexplored. Through the execution of a single Python script, we present a fully automated pipeline leveraging the Claude API to generate systematic reviews from literature search through manuscript completion without human intervention. Our pipeline processes hundreds of papers through iterative API calls for inclusion evaluation, information extraction, and synthesis, achieving citation accuracy rates of 95.87% through controlled text-restriction strategies that mitigate hallucination. In a blinded evaluation, six board-certified hematopathologists rated AI-generated systematic reviews (mean quality score: 3.4-3.66/5) higher than a published human-authored review (2.6/5) on the same topic, yet failed to reliably distinguish AI from human authorship. Notably, the human-written review was most frequently misidentified as AI-generated, revealing systematic biases in expert perception of AI capabilities. While demonstrating superior prose quality and topic coherence, AI-generated reviews exhibited increased repetition and had to be restricted to only referencing a select number of papers per section, highlighting fundamental trade-offs between automation scale and information breadth. Our findings establish both the technical feasibility and critical limitations of LLM-driven knowledge synthesis, raising urgent questions about verification standards, disclosure practices, and potential misuse in academic publishing. As automated high-quality scientific writing becomes computationally trivial, we argue for establishing transparent integration frameworks and enhanced AI literacy among domain experts to preserve scientific integrity while harnessing efficiency gains.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 95%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 95%
- Transparency in peer review: Exploring the content and tone of reviewers' confidential comments to editors 95%
Similar papers in this journal
- Network Graph Representation of COVID-19 Scientific Publications to Aid Knowledge Discovery 94%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.