How good are large language models for automated data extraction from randomized trials?
Sun, Z.; Zhang, R.; Doi, S. A.; Furuya-Kanamori, L.; Yu, T.; Lin, L.; Xu, C.
Show abstract
In evidence synthesis, data extraction is a crucial procedure, but it is time intensive and prone to human error. The rise of large language models (LLMs) in the field of artificial intelligence (AI) offers a solution to these problems through automation. In this case study, we evaluated the performance of two prominent LLM-based AI tools for use in automated data extraction. Randomized trials from two systematic reviews were used as part of the case study. Prompts related to each data extraction task (e.g., extract event counts of control group) were formulated separately for binary and continuous outcomes. The percentage of correct responses (Pcorr) was tested in 39 randomized controlled trials reporting 10 binary outcomes and 49 randomized controlled trials reporting one continuous outcome. The Pcorr and agreement across three runs for data extracted by two AI tools were compared with well-verified metadata. For the extraction of binary events in the treatment group across 10 outcomes, the Pcorr ranged from 40% to 87% and from 46% to 97% for ChatPDF and for Claude, respectively. For continuous outcomes, the Pcorr ranged from 33% to 39% across six tasks (Claude only). The agreement of the response between the three runs of each task was generally good, with Cohens kappa statistic ranging from 0.78 to 0.96 and from 0.65 to 0.82 for ChatPDF and Claude, respectively. Our results highlight the potential of ChatPDF and Claude for automated data extraction. Whilst promising, the percentage of correct responses is still unsatisfactory and therefore substantial improvements are needed for current AI tools to be adopted in research practice. Highlights1. What is already knownO_LIIn evidence synthesis, data extraction is a crucial procedure, but it is time intensive and prone to human error, with reported data extraction error rates at meta-analyses level reaching up to 67%. C_LIO_LIThe rise of large language models (LLMs) in the field of artificial intelligence (AI) offers a solution to these problems through automation. C_LI 2. What is newO_LIIn this case study, we investigated the performance of two AI tools for data extraction and confirmed that AI tools can reach the same or better performance than humans in terms of data extraction from randomized trials for binary outcomes. C_LIO_LIHowever, AI tools performed poorly at extracting data from continuous outcomes. C_LI 3. Potential impact for Research Synthesis Methods readers outside the authors fieldO_LIOur study suggests LLMs have great potential in assisting data extraction in evidence syntheses through (semi-)automation. Further efforts are needed to improve accuracy, especially for continuous outcomes data. C_LI
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 97%
- Development of a search filter to retrieve reports of interrupted time series studies from MEDLINE and PubMed 96%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 96%
Similar papers in this journal
- Incorporating Preprints in Systematic Reviews: A Preliminary Study of a Novel Method for Rapid Evidence Synthesis 95%
- A Web-based Tool for Automatically linking Clinical Trials to their Publications 94%
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 93%
Similar papers in this journal
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 95%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 95%
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 95%
Similar papers in this journal
- Investigator-initiated versus industry-sponsored trials – Visibility and relevance of randomized controlled trials in clinical practice guidelines (IMPACT) 95%
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 94%
- Completeness of reporting of clinical prediction models developed using supervised machine learning: A systematic review 94%
Similar papers in this journal
- GPT for RCTs?: Using AI to measure adherence to reporting guidelines 95%
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 95%
- The incubation period of COVID-19: A rapid systematic review and meta-analysis of observational research 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.