AI-Assisted Data Extraction with a Large Language Model: A Study Within Reviews
Gartlehner, G.; Kugley, S.; Crotty, K.; Viswanathan, M.; Dobrescu, A.; Nussbaumer-Streit, B.; Booth, G.; Treadwell, J.; Han, J. M.; Wagner, J.; Apaydin, E.; Coppola, E.; Maglione, M.; Hilscher, R.; Chew, R.; Pilar, M.; Swanton, B.; Kahwati, L.
Show abstract
BackgroundData extraction is a critical but error-prone and labor-intensive task in evidence synthesis. Unlike other artificial intelligence (AI) technologies, large language models (LLMs) do not require labeled training data for data extraction. ObjectiveTo compare an AI-assisted to a traditional y data extraction process. DesignStudy within reviews (SWAR) utilizing a prospective, parallel group comparison with blinded data adjudicators. SettingWorkflow validation within six ongoing systematic reviews of interventions under real-world conditions. InterventionInitial data extraction using an LLM (Claude versions 2.1, 3.0 Opus, and 3.5 Sonnet) verified by a human reviewer. MeasurementsConcordance, time on task, accuracy, recall, precision, and error analysis. ResultsThe six systematic reviews of the SWAR contributed 9,341 data elements, extracted from 63 studies. Concordance between the two methods was 77.2%. The accuracy of the AI-assisted approach compared with enhanced human data extraction was 91.0%, with a recall of 89.4% and a precision of 98.9%. The AI-assisted approach had fewer incorrect extractions (9.0% vs. 11.0%) and similar risks of major errors (2.5% vs. 2.7%) compared to the traditional human-only method, with a median time saving of 41 minutes per study. Missed data items were the most frequent errors in both approaches. LimitationsAssessing the concordance of data extractions and classifying errors required subjective judgment. Tracking time on task consistently was challenging. ConclusionThe use of an LLM can improve accuracy of data extraction and save time in evidence synthesis. Results reinforce previous findings that human-only data extraction is prone to errors. Primary Funding SourceUS Agency for Healthcare Research and Quality, RTI International RegistrationSWAR28 Gerald Gartlehner (2023 FEB 11 2102).pdf
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Use of estimands in cluster randomised trials: a review 92%
- A modular pipeline for natural language processing-screened human abstraction of a pragmatic trial outcome from electronic health records 91%
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 90%
Similar papers in this journal
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 97%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 96%
- Exploring the potential of Claude 2 for risk of bias assessment: Using a large language model to assess randomized controlled trials with RoB 2 96%
Similar papers in this journal
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 96%
- Updating the PRISMA reporting guideline for network meta-analysis: a scoping review 96%
- Characteristics and completeness of reporting of systematic reviews of prevalence studies in adult populations: a meta-epidemiological study 95%
Similar papers in this journal
- Protocol for the development of a tool (INSPECT-SR) to identify problematic randomised controlled trials in systematic reviews of health interventions 96%
- Agreeability testing of AMSTAR-PF, a tool for quality appraisal of systematic reviews of prognostic factor studies 96%
- Comparison of preprints and final journal publications from COVID-19 Studies: Discrepancies in results reporting and spin in interpretation 96%
Similar papers in this journal
- COVID-19-related research data availability and quality according to the FAIR principles: A meta-research study 94%
- Exploring scalable assessment methods for terminated trials in ClinicalTrials.gov: A cohort analysis of German and Californian trials 94%
- Transparency in peer review: Exploring the content and tone of reviewers' confidential comments to editors 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.