MechaScreener: Large Language Model-Based Automated Screening for Systematic Reviews and Research
Forbes, C.; Carter, M.; Hudson, C.; Glasziou, P.; Clark, J.
Show abstract
Systematic Reviews (SRs) are the gold standard for evidence synthesis, but the manual title and abstract screening of thousands of references creates a severe bottleneck. Existing automated tools have historically struggled to achieve the near-perfect recall (sensitivity) required for reliable reviews. We developed MechaScreener as a "zero-shot" automated screening tool that utilises a Large Language Model (LLM) to rank article relevance. The tool requires no initial training data or manual pre-screening, as MechaScreener directly applies user-provided question elements (PICO) or inclusion/exclusion criteria to assign an inclusion probability score (1-5) to each reference. We evaluated the tool in two phases: a development phase using five reference libraries to optimise prompts, and an independent evaluation phase using 10 diverse Cochrane review libraries (comprising both randomised controlled trials and non-RCTs) containing over 58,000 references. In the evaluation dataset, MechaScreener achieved a perfect mean recall of 1.00 (100%, pooled 95% CI: 0.98-1.00), ensuring no relevant articles were missed. Concurrently, it achieved an overall mean specificity of 0.61 (61%, pooled 95% CI: 0.59-0.60). Specificity varied: from 0.21 in broad public health topics to 0.91 in precise pharmacological interventions-reflecting the tools built-in conservatism when evaluating ambiguous abstracts. By safely eliminating over 60% of irrelevant literature during the initial screening phase without compromising recall, MechaScreener functions as a highly reliable but low-effort "first-pass" filter, allowing researchers to substantially reduce manual workloads and reallocate resources toward full-text review and data extraction.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Catchii: empowering literature review screening in healthcare 97%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 96%
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 96%
Similar papers in this journal
- Incorporating Preprints in Systematic Reviews: A Preliminary Study of a Novel Method for Rapid Evidence Synthesis 96%
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 94%
- A Novel Question-Answering Framework for Automated Abstract Screening Using Large Language Models 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 94%
- Modelling the impact of behavioural interventions during pandemics: A systematic review 93%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.