Performance of Large Language Models in Automated Medical Literature Screening: A Systematic Review and Meta-analysis
Chenggong, X.; Weichang, K.; Liuting, P.; Diaoxin, Q.; Yuxuan, Y.; Bin, W.; Liang, H.
Show abstract
ObjectiveTo systematically evaluate the diagnostic performance of large language models (LLMs) in automated medical literature screening and to determine their potential role in supporting evidence synthesis workflows. MethodsA systematic review and meta-analysis was conducted according to PRISMA DTA guidance. PubMed, Web of Science, Embase, the Cochrane Library and Google Scholar were searched from 1 January 2022 to 17 November 2025. Studies assessing LLMs for automated title and abstract screening or full-text eligibility assessment in medical literature were included. Diagnostic accuracy metrics were extracted and pooled using a bivariate random effects model and hierarchical summary receiver operating characteristic (HSROC) analysis. Subgroup analyses and meta-regression were performed to explore sources of heterogeneity. ResultsEighteen studies published between 2023 and 2025 were included. In title and abstract screening, the pooled sensitivity was 0.92 and pooled specificity was 0.94. The SROC area under the curve (AUC) reached 0.98. In full-text screening, pooled sensitivity and specificity both reached 0.99 and the AUC was 0.99. Prompt strategies incorporating examples or chain-of-thought reasoning significantly improved sensitivity. Across studies, most models were deployed without task specific fine tuning and still achieved strong performance. Subgroup analyses and meta regression did not identify significant sources of heterogeneity. Many studies also reported substantial efficiency gains, including large reductions in screening workload, time and cost. ConclusionLLMs demonstrate high diagnostic accuracy for automated medical literature screening, particularly in full-text assessment. These models show strong potential as high sensitivity assistive tools that can substantially reduce manual screening burden while supporting evidence synthesis. Further methodological optimization and validation in large scale real-world settings are required to establish their long term role in evidence-based medicine.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning for identifying relevant publications in updates of systematic reviews of diagnostic test studies 96%
- Citation tracking for systematic literature searching: a scoping review 95%
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 95%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 96%
- Characteristics and completeness of reporting of systematic reviews of prevalence studies in adult populations: a meta-epidemiological study 94%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 94%
Similar papers in this journal
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 94%
- Adoption of the OMOP CDM for Cancer Research using Real-world Data: Current Status and Opportunities 93%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 95%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 93%
- Modelling the impact of behavioural interventions during pandemics: A systematic review 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.