Large language models for abstract screening in systematic- and scoping reviews: A diagnostic test accuracy study
Krag, C. H.; Balschmidt, T.; Bruun, F.; Brejnebol, M. W.; Xu, J. J.; Boesen, M.; Andersen, M. B.; Müller, F. C.
Show abstract
IntroductionWe investigated if large language models (LLMs) can be used for abstract screening in systematic- and scoping reviews. MethodsTwo broad reviews were designed: a systematic review structured according to the PRISMA guideline with abstract inclusion based on PICO criteria; and a scoping review, where we defined abstract characteristics and features of interest to look for. For both reviews 500 abstracts were sampled. Two readers independently screened abstracts with disagreements handled with arbitrations or consensus, which served as the reference standard. The abstracts were analysed by six LLMs (GPT-4o, GPT-4T, GPT-3.5, Claude3-Opus, Claude3-Sonnet, and Claude3-Haiku). Primary outcomes were diagnostic test accuracy measures for abstract inclusion, abstract characterisation and feature of interest detection. Secondary outcome was the degree of automation using LLMs as a function of the error rate. ResultsIn the systematic review 12 studies were marked as include by the human consensus. GPT-4o, GPT-4T, and Claude3-Opus achieved the highest accuracies (97% to 98%) comparable to the human readers (96% and 98%), although sensitivity was low (33% to 50%). In the scoping review 130 features of interest were present and the LLMs achieved sensitivities between 74-84%, comparable to the human readers (73% and 86%). The specificity of GPT-4o (98%) and GPT-4T (>99%) greatly surpassed the other LLMs (between 33% and 93%). For abstract characterization all LLMs achieved above 95% accuracy for language, manuscript type and study participant characterisation. For characterisation of disease-specific features only GPT-4T and GPT-4o showed very high accuracy. For abstract inclusion the highest automation rate (91%) at the lowest error rate (8%) was achieved by use of two LLMs with disagreement solved by human arbitration. An LLM pre screening before human abstract screening achieved an automation rate of 55% with no missed abstracts. ConclusionAbstract characterisation and specific feature of interest detection with LLMs is feasible and accurate with GPT-4o and GPT-4T. The majority of abstract screenings for systematic reviews can be automated with use of LLMs, at low error rates.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 97%
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 97%
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 96%
Similar papers in this journal
Similar papers in this journal
- Agreeability testing of AMSTAR-PF, a tool for quality appraisal of systematic reviews of prognostic factor studies 96%
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 95%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 95%
Similar papers in this journal
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 94%
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 93%
- Modelling the impact of behavioural interventions during pandemics: A systematic review 93%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 96%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 95%
- Characteristics and completeness of reporting of systematic reviews of prevalence studies in adult populations: a meta-epidemiological study 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.