Making Broad Evidence Synthesis Feasible: An LLM Screening Agent for Meta-Analyses Applied To Suicide Prevention
Dobin, D.; Witmer, A. M.; Sweeney, F.; Ryan, T.; Cimino, A.; Haroz, E. E.; Nestadt, P. S.; Wilcox, H. C.
Show abstract
Importance. Systematic reviews and meta-analyses inform suicide-prevention policy and practice, but broad database searches are difficult to screen manually. This limits capture of upstream interventions, such as economic policies, with indirect effects on suicide. Reliable automated screening could make broader and more comprehensive evidence syntheses feasible. Objective. To develop and validate ScreenAgent, a large language model (LLM) agent for title and abstract screening, and a review-specific method for prospectively estimating screening performance. Design, Setting, and Participants. ScreenAgent was validated internally on a prospective meta-analysis, and externally on two published systematic reviews. The correct include and exclude decisions followed standard systematic-review screening methodology. Exposures. ScreenAgent, an LLM agent returning structured include-or-exclude decisions. Records it marked for inclusion were re-checked by a second, cascade pass using a higher-effort LLM. For the external reviews, the agent's prompt was tuned automatically on a small set of labeled examples. Main Outcomes and Measures. We calculated sensitivity, specificity, workload reduction (the percentage of records removed from human review), and agent-versus-human reliability via Cohen kappa. Sensitivity was estimated by direct comparison (internal) and 5-fold cross-validation (external). Results. In the internal validation, ScreenAgent identified 43 of 44 eligible studies (sensitivity 97.7%; 95% CI, 88.2%-99.6%) with a generic prompt applied without any review-specific optimization, specificity 98.0%, and a measured full-corpus workload reduction of 99.4%. The cost was $855.91 for the full 201,064-record corpus (0.43 US cents per record). Agent-versus-human-consensus agreement exceeded human-versus-human agreement (Cohen kappa 0.75 vs 0.64; percent agreement 97.3% vs 95.4%). For two external validation studies, automatic tuning resulted in a cross-validated sensitivity of 95.9% (95% CI, 90.0%-98.4%) and 97.4% (90.9%-99.3%), with workload reductions of 97.4% and 98.4%. Conclusions and Relevance. Suicide prevention efforts often require rapid consolidation of evidence because of the inherent challenges of single studies trying to prevent rare outcomes. On both internal and external validation sets, ScreenAgent identified nearly all eligible studies with human-level reliability for a fraction of a US cent per record while keeping human reviewers as the final arbiters. By making broad searches feasible and screening performance measurable beforehand, this approach can serve as a transparent methodology to strengthen the speed at which we can inform and advance suicide prevention efforts.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- To Include or Not to Include? A prescription from the pharmacy on how to use active learning assisted screening in systematic reviews 92%
- Repurposing Existing Medications for Coronavirus Disease 2019: Protocol for a Rapid and Living Systematic Review 92%
- Evidence-Based, Cost-Effective Interventions To Suppress The COVID-19 Pandemic: A Systematic Review 91%
Similar papers in this journal
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 93%
- Individual Participant Data Network Meta-analysis of psychosocial interventions for survivors of intimate partner violence: Study protocol 92%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 91%
Similar papers in this journal
- A Systematic Review of Machine Learning-based Prognostic Models for Acute Pancreatitis: Towards Improving Methods and Reporting Quality 90%
- Accuracy and clinical effectiveness of risk prediction tools for pressure injury occurrence: An umbrella review 90%
- Asymptomatic SARS-CoV-2 infections: a living systematic review and meta-analysis 89%
Similar papers in this journal
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 95%
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 94%
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 94%
Similar papers in this journal
- Agreeability testing of AMSTAR-PF, a tool for quality appraisal of systematic reviews of prognostic factor studies 94%
- The incubation period of COVID-19: A rapid systematic review and meta-analysis of observational research 93%
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.