Accelerating the pace and accuracy of systematic reviews using AI: a validation study
Zhan, J.; Suvada, K.; Xu, M.; Tian, W.; Cara, K. C.; Wallace, T. C.; Ali, M. K.
Show abstract
BackgroundArtificial intelligence (AI) can greatly enhance efficiency in systematic literature reviews and meta-analyses, but its accuracy in screening titles/abstracts and full-text articles is uncertain. ObjectivesThis study evaluated the performance metrics (sensitivity, specificity) of a GPT-4 AI program, Review Copilot, against human decisions (gold standard) in screening titles/abstracts and full-text articles from four published systematic reviews/meta-analyses. Research DesignParticipant data from four already-published systematic literature reviews were used for this validation study. This was a study comparing Review Copilot to human decision-making (gold standard) in screening titles/abstracts and full-text articles for systematic reviews/meta-analyses. The four studies that were used in this study included observational studies and randomized control trials. Review Copilot operates on the OpenAI, GPT-4 server. We examined the performance metrics of Review Copilot to include and exclude titles/abstracts and full-text articles as compared to human decisions in four systematic reviews/meta-analyses. Sensitivity, specificity, and balanced accuracy of title/abstract and full-text screening were compared between Review Copilot and human decisions. ResultsReview Copilots sensitivity and specificity for title/abstract screening were 99.2% and 83.6%, respectively, and 97.6% and 47.4% for full-text screening. The average agreement between two runs was 95.4%, with a kappa statistic of 0.83. Review Copilot screened in one-quarter of the time compared to humans. ConclusionsAI use in systematic reviews and meta-analyses is inevitable. Health researchers must understand these technologies strengths and limitations to ethically leverage them for research efficiency and evidence-based decision-making in health.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Citation tracking for systematic literature searching: a scoping review 97%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 97%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 97%
Similar papers in this journal
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 97%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 96%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 96%
Similar papers in this journal
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 96%
- Comparison of preprints and final journal publications from COVID-19 Studies: Discrepancies in results reporting and spin in interpretation 96%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 96%
Similar papers in this journal
- To Include or Not to Include? A prescription from the pharmacy on how to use active learning assisted screening in systematic reviews 95%
- Repurposing Existing Medications for Coronavirus Disease 2019: Protocol for a Rapid and Living Systematic Review 93%
- Evidence-Based, Cost-Effective Interventions To Suppress The COVID-19 Pandemic: A Systematic Review 93%
Similar papers in this journal
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 96%
- Does advance contact with research participants increase response to questionnaires: A Systematic Review and meta-Analysis 95%
- Does pre-notification increase questionnaire response rates: a nested randomised control trial 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.