A PRISMA-Aligned Agentic Framework for Medical Systematic Reviews and Evidence Synthesis
Huang, H.; Zheng, Q.; Qiu, P.; Zhao, W.; Zhang, Y.; Xie, W.; Wang, Y.; Zhang, X.; Wu, C.
Show abstract
Medical systematic reviews are central to evidence-based medicine, but they remain slow, labor-intensive, and difficult to maintain under the full Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) workflow. Recent LLM-based deep research agents offer a promising route to addressing this challenge, yet reliable deployment in medical systematic reviews remains limited by insufficient clinical domain knowledge and inconsistent adherence to evidence-based methodological standards across the full workflow. We address these gaps with MedSR-Copilot, a PRISMA-aligned multi-agent copilot that decomposes review automation into literature retrieval, coarse-to-fine screening, data extraction, Risk-of-Bias assessment, and evidence synthesis, while preserving structured intermediate artifacts throughout the workflow. We further introduce MedSR-Bench, an end-to-end benchmark for evaluating systems beyond isolated subtasks, from review input to final evidence-synthesis conclusions. MedSR-Copilot completes medical systematic reviews end-to-end under the full PRISMA workflow, achieving 63.6% human-aligned conclusions, 18.3 percentage points above the best baseline among strong general-purpose LLMs and prior automated review systems. In a human-AI collaboration study involving 23 analysis groups across four systematic review topics, MedSR-Copilot, used as a copilot, reduces end-to-end review time by 64.9% and improves final conclusion accuracy by 27.4 percentage points compared with routine-practice workflows. Together, these results demonstrate the reliability and efficiency of MedSR-Copilot as a medical research copilot and suggest a practical path toward trustworthy review automation.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Comparing scientific abstracts generated by ChatGPT to original abstracts using an artificial intelligence output detector, plagiarism detector, and blinded human reviewers 92%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 91%
Similar papers in this journal
- Machine learning for identifying relevant publications in updates of systematic reviews of diagnostic test studies 93%
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 93%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 92%
Similar papers in this journal
- A modular pipeline for natural language processing-screened human abstraction of a pragmatic trial outcome from electronic health records 95%
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 91%
- Use of estimands in cluster randomised trials: a review 89%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 92%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 92%
- Re-use of trial data in the first 10 years of the data-sharing policy of the Annals of Internal Medicine: a survey of published studies 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.