Back

Fully Automated Systematic Review Generation via Large Language Models: Quality Assessment and Implications for Scientific Publishing

McLaughlin, L.; Walz, M. S.; Arries, C.

2026-02-23 health informatics
10.64898/2026.02.18.26346559 medRxiv
Show abstract

Large language models (LLMs) are increasingly transforming scientific workflows, yet their application to rigorous evidence synthesis remains underexplored. Through the execution of a single Python script, we present a fully automated pipeline leveraging the Claude API to generate systematic reviews from literature search through manuscript completion without human intervention. Our pipeline processes hundreds of papers through iterative API calls for inclusion evaluation, information extraction, and synthesis, achieving citation accuracy rates of 95.87% through controlled text-restriction strategies that mitigate hallucination. In a blinded evaluation, six board-certified hematopathologists rated AI-generated systematic reviews (mean quality score: 3.4-3.66/5) higher than a published human-authored review (2.6/5) on the same topic, yet failed to reliably distinguish AI from human authorship. Notably, the human-written review was most frequently misidentified as AI-generated, revealing systematic biases in expert perception of AI capabilities. While demonstrating superior prose quality and topic coherence, AI-generated reviews exhibited increased repetition and had to be restricted to only referencing a select number of papers per section, highlighting fundamental trade-offs between automation scale and information breadth. Our findings establish both the technical feasibility and critical limitations of LLM-driven knowledge synthesis, raising urgent questions about verification standards, disclosure practices, and potential misuse in academic publishing. As automated high-quality scientific writing becomes computationally trivial, we argue for establishing transparent integration frameworks and enhanced AI literacy among domain experts to preserve scientific integrity while harnessing efficiency gains.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.