Comparison of Elicit AI and Traditional Literature Searching in Systematic Reviews using Four Case Studies
Golder, S.; Lau, O.
Show abstract
BackgroundElicit AI aims to simplify and accelerate the systematic review process without compromising accuracy. However, research on Elicits performance is limited. ObjectivesTo determine whether Elicit AI is a viable tool for systematic literature searches. MethodsWe compared the included studies in four systematic reviews to those identified searching with Elicit. We calculated sensitivity, precision and observed patterns in the performance of Elicit. ResultsElicit had an average of 39.6% precision (26.7% - 46.2%) which was higher than the 7.55% average of the original reviews (0.65% - 14.7%). However, the sensitivity of Elicit was poor, averaging 37.9% (25.5% - 69.2%) compared to 93.5% (87.2% - 98.0%) in the original reviews. Elicit also identified some included studies not identified by the original searches. DiscussionAt the time of this evaluation, Elicit did not search with high enough sensitivity to replace traditional literature searching. However, the high precision of searching in Elicit could prove useful for preliminary searches, and the unique studies identified mean that Elicit can be used by researchers as a useful adjunct. ConclusionWhilst Elicit searches are currently not sensitive enough to replace traditional searching, Elicit is continually improving, and further evaluations should be undertaken as new developments take place. Key MessagesO_LIAI tools, such as Elicit, have been developed to improve the efficiency of systematic review processes, including the identification of studies. C_LIO_LIUsing four case study systematic reviews Elicit searches had a sensitivity between 25.5% and 69.2% (37.9% average) and precision between 26.7% and 46.2% (39.6% average). C_LIO_LIElicit identified some unique studies that met the inclusion criteria for each of the case study systematic reviews. C_LIO_LIElicit is constantly improving and developing its systems, thus independent researchers should continue to evaluate its performance. C_LI
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- From smartphone data to clinically relevant predictions: A systematic review of digital phenotyping methods in depression 88%
- The neuroprotective effects of estrogen and estrogenic compounds in spinal cord injury 87%
- Problematic usage of the internet and eating disorders: a multifaceted, systematic review and meta-analysis 87%
Similar papers in this journal
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 95%
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 94%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 94%
Similar papers in this journal
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 94%
- Strategies used to manage overlap of primary study data by exercise-related overviews. Protocol for a systematic methodological review 94%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 93%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 94%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 93%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 93%
Similar papers in this journal
- Does advance contact with research participants increase response to questionnaires: A Systematic Review and meta-Analysis 96%
- Does pre-notification increase questionnaire response rates: a nested randomised control trial 95%
- Quantitative bias analysis for mismeasured variables in health research: a review of software tools 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.