Comparison of large language models for citation screening: A protocol for a prospective study
Oami, T.; Okada, Y.; Nakada, T.-a.
Show abstract
BackgroundSystematic reviews require labor-intensive and time-consuming processes. Large language models (LLMs) have been recognized as promising tools for citation screening; however, the performance of LLMs in screening citations remained to be determined yet. This study aims to evaluate the potential of three leading LLMs - GPT-4o, Gemini 1.5 Pro, and Claude 3.5 Sonnet for literature screening. MethodsWe will conduct a prospective study comparing the accuracy, efficiency, and cost of literature citation screening using the three LLMs. Each model will perform literature searches for predetermined clinical questions from the Japanese Clinical Practice Guidelines for Management of Sepsis and Septic Shock (J-SSCG). We will measure and compare the time required for citation screening using each method. The sensitivity and specificity of the results from the conventional approach and each LLM-assisted process will be calculated and compared. Additionally, we will assess the total time spent and associated costs for each method to evaluate workload reduction and economic efficiency. Trial registrationThis research is submitted with the University hospital medical information network clinical trial registry (UMIN-CTR) [UMIN000054783].
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 95%
- Introducing the EMPIRE Index: A novel, value-based metric framework to measure the impact of medical publications 94%
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 94%
Similar papers in this journal
- Catchii: empowering literature review screening in healthcare 96%
- Fast and frugal decision tree for the rapid critical appraisal of systematic reviews 94%
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 93%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 94%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 93%
- The impact of retracted randomised controlled trials on systematic reviews and clinical practice guidelines: a meta-epidemiological study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.