Back

A PRISMA-Aligned Agentic Framework for Medical Systematic Reviews and Evidence Synthesis

Huang, H.; Zheng, Q.; Qiu, P.; Zhao, W.; Zhang, Y.; Xie, W.; Wang, Y.; Zhang, X.; Wu, C.

2026-08-02 health informatics
10.64898/2026.07.30.26359375 medRxiv
Show abstract

Medical systematic reviews are central to evidence-based medicine, but they remain slow, labor-intensive, and difficult to maintain under the full Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) workflow. Recent LLM-based deep research agents offer a promising route to addressing this challenge, yet reliable deployment in medical systematic reviews remains limited by insufficient clinical domain knowledge and inconsistent adherence to evidence-based methodological standards across the full workflow. We address these gaps with MedSR-Copilot, a PRISMA-aligned multi-agent copilot that decomposes review automation into literature retrieval, coarse-to-fine screening, data extraction, Risk-of-Bias assessment, and evidence synthesis, while preserving structured intermediate artifacts throughout the workflow. We further introduce MedSR-Bench, an end-to-end benchmark for evaluating systems beyond isolated subtasks, from review input to final evidence-synthesis conclusions. MedSR-Copilot completes medical systematic reviews end-to-end under the full PRISMA workflow, achieving 63.6% human-aligned conclusions, 18.3 percentage points above the best baseline among strong general-purpose LLMs and prior automated review systems. In a human-AI collaboration study involving 23 analysis groups across four systematic review topics, MedSR-Copilot, used as a copilot, reduces end-to-end review time by 64.9% and improves final conclusion accuracy by 27.4 percentage points compared with routine-practice workflows. Together, these results demonstrate the reliability and efficiency of MedSR-Copilot as a medical research copilot and suggest a practical path toward trustworthy review automation.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.