Back

Comparison of Elicit AI and Traditional Literature Searching in Systematic Reviews using Four Case Studies

Golder, S.; Lau, O.

2025-06-27 health informatics
10.1101/2025.06.17.25329772 medRxiv
Show abstract

BackgroundElicit AI aims to simplify and accelerate the systematic review process without compromising accuracy. However, research on Elicits performance is limited. ObjectivesTo determine whether Elicit AI is a viable tool for systematic literature searches. MethodsWe compared the included studies in four systematic reviews to those identified searching with Elicit. We calculated sensitivity, precision and observed patterns in the performance of Elicit. ResultsElicit had an average of 39.6% precision (26.7% - 46.2%) which was higher than the 7.55% average of the original reviews (0.65% - 14.7%). However, the sensitivity of Elicit was poor, averaging 37.9% (25.5% - 69.2%) compared to 93.5% (87.2% - 98.0%) in the original reviews. Elicit also identified some included studies not identified by the original searches. DiscussionAt the time of this evaluation, Elicit did not search with high enough sensitivity to replace traditional literature searching. However, the high precision of searching in Elicit could prove useful for preliminary searches, and the unique studies identified mean that Elicit can be used by researchers as a useful adjunct. ConclusionWhilst Elicit searches are currently not sensitive enough to replace traditional searching, Elicit is continually improving, and further evaluations should be undertaken as new developments take place. Key MessagesO_LIAI tools, such as Elicit, have been developed to improve the efficiency of systematic review processes, including the identification of studies. C_LIO_LIUsing four case study systematic reviews Elicit searches had a sensitivity between 25.5% and 69.2% (37.9% average) and precision between 26.7% and 46.2% (39.6% average). C_LIO_LIElicit identified some unique studies that met the inclusion criteria for each of the case study systematic reviews. C_LIO_LIElicit is constantly improving and developing its systems, thus independent researchers should continue to evaluate its performance. C_LI

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.