Fast clinical trial identification using fuzzy-search elastic searches: retrospective validation with high-quality Cochrane benchmark
Otte, W. M.; van IJzendoorn, D. G.; Habets, P. C.; Vinkers, C. H.
Show abstract
The synthesis of treatment effects relies on systematic reviews of intervention trials. This process is often laborious due to the need for precise search queries and manual study identification. Recent advancements in database architecture and natural language processing (NLP) offer a potential solution by enabling faster, more flexible searches and automated extraction of information from unstructured texts. Our study assesses the effectiveness of NLP-based literature searches within a novel database structure in comparison to the Cochrane Database of Systematic Reviews. We created a user-friendly elastic search database containing 36 million PubMed-indexed entries. We developed reliable filters for identifying randomized clinical trials and clinical intervention studies, as well as extracting relevant subtext related to population and intervention. Our results indicate a high precision of 0.74, recall of 0.81, and F1-score of 0.77 for population subtext, and a precision of 0.70, recall of 0.71, and an F1-score of 0.70 for intervention subtext. Our approach efficiently identified included studies in 90% of systematic reviews, missing no more than two trials compared to Cochrane. Furthermore, it produced fewer total hits than a comparable PubMed keyword search, demonstrating the potential of the new database structure to enhance the efficiency and effectiveness of aggregating clinical evidence.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 97%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 96%
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 96%
Similar papers in this journal
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 96%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 96%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 95%
Similar papers in this journal
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 92%
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 90%
- Pregnancy pharmacoepidemiology: How often are key methodological elements reported in publications? 90%
Similar papers in this journal
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 95%
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 95%
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 94%
Similar papers in this journal
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 96%
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 95%
- Reporting of Retrospective Registration in Clinical Trial Publications 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.