Sensitivity, specificity and avoidable workload of using a large language models for title and abstract screening in systematic reviews and meta-analyses
Tran, V.-T.; Gartlehner, G.; Yaacoub, S.; Boutron, I.; Schwingshackl, L.; Stadelmaier, J.; Sommer, I.; Aboulayeh, F.; Afach, S.; Meerpohl, J.; Ravaud, P.
Show abstract
ImportanceSystematic reviews are time-consuming and are still performed predominately manually by researchers despite the exponential growth of scientific literature. ObjectiveTo investigate the sensitivity, specificity and estimate the avoidable workload when using an AI-based large language model (LLM) (Generative Pre-trained Transformer [GPT] version 3.5-Turbo from OpenAI) to perform title and abstract screening in systematic reviews. Data SourcesUnannotated bibliographic databases from five systematic reviews conducted by researchers from Cochrane Austria, Germany and France, all published after January 2022 and hence not in the training data set from GPT 3.5-Turbo. DesignWe developed a set of prompts for GPT models aimed at mimicking the process of title and abstract screening by human researchers. We compared recommendations from LLM to rule out citations based on title and abstract with decisions from authors, with a systematic reappraisal of all discrepancies between LLM and their original decisions. We used bivariate models for meta-analyses of diagnostic accuracy to estimate pooled estimates of sensitivity and specificity. We performed a simulation to assess the avoidable workload from limiting human screening on title and abstract to citations which were not "ruled out" by the LLM in a random sample of 100 systematic reviews published between 01/07/2022 and 31/12/2022. We extrapolated estimates of avoidable workload for health-related systematic reviews assessing therapeutic interventions in humans published per year. ResultsPerformance of GPT models was tested across 22,666 citations. Pooled estimates of sensitivity and specificity were 97.1% (95%CI 89.6% to 99.2%) and 37.7%, (95%CI 18.4% to 61.9%), respectively. In 2022, we estimated the workload of title and abstract screening for systematic reviews to range from 211,013 to 422,025 person-hours. Limiting human screening to citations which were not "ruled out" by GPT models could reduce workload by 65% and save up from 106,268 to 276,053-person work hours (i.e.,66 to 172-person years of work), every year. Conclusions and RelevanceAI systems based on large language models provide highly sensitive and moderately specific recommendations to rule out citations during title and abstract screening in systematic reviews. Their use to "triage" citations before human assessment could reduce the workload of evidence synthesis.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 97%
- Investigation of bias due to selective inclusion of study effect estimates in meta-analyses of nutrition research 97%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 97%
Similar papers in this journal
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 97%
- Methods used to select results to include in meta-analyses of nutrition research: a meta-research study 96%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 95%
Similar papers in this journal
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 95%
- Does advance contact with research participants increase response to questionnaires: A Systematic Review and meta-Analysis 95%
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 94%
Similar papers in this journal
- Tool to assess risk of bias due to missing evidence in network meta-analysis (ROB-MEN): elaboration and examples 96%
- Transparency and reporting characteristics of COVID-19 randomized controlled trials 94%
- Evidence of unexplained discrepancies between planned and conducted statistical analyses: a review of randomized trials 93%
Similar papers in this journal
- To Include or Not to Include? A prescription from the pharmacy on how to use active learning assisted screening in systematic reviews 94%
- Evidence-Based, Cost-Effective Interventions To Suppress The COVID-19 Pandemic: A Systematic Review 94%
- Repurposing Existing Medications for Coronavirus Disease 2019: Protocol for a Rapid and Living Systematic Review 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.