Automating clinical trial outcome identification and misreporting detection using RegCheck
Cummins, J.; Drysdale, H.; Elson, M.; Hussey, I.; Goldacre, B.; DeVito, N. J.
Show abstract
Objective To evaluate the accuracy and cost of RegCheck, an automated large language model (LLM)-based workflow, for identifying clinical trial outcomes and detecting outcome misreporting by comparing its outputs with manual assessments from the COMPare Trials project. Design Validation study. Setting Sixty-two clinical trials originally assessed in the COMPare Trials project, sampled from five high impact general medical journals. Participants Published clinical trial reports and their corresponding prespecified registrations and/or protocols. Main outcome measures Four prespecified research questions were examined. RQ1 assessed outcome extraction recall relative to COMPare. RQ2 assessed accuracy of outcome classification as primary, secondary, or non-prespecified. RQ3 assessed accuracy of misreporting detection relative to COMPare, with additional manual adjudication of discrepancies between RegCheck and COMPare. RQ4 assessed the average per-paper cost of running the automated workflow. Results Across the validation papers, RegCheck achieved 91.2% outcome extraction recall relative to COMPare, and 83.6% outcome classification accuracy. For detection of outcome misreporting, RegCheck's overall accuracy was 85.6%. However, after resolving discrepancies with the original human judgements (which frequently favoured RegCheck's judgement), revised accuracy for outcome misreporting detection was 94.8%. The mean cost of running the workflow was 5.94 USD per paper. Conclusions RegCheck achieved high overall performance with a rigorous manual benchmark for identifying prespecified and reported outcomes in clinical trials, and detecting outcome misreporting, while operating at very low marginal cost. Adjudication of discrepant judgements suggested that RegCheck frequently identified valid issues not captured in the reference standard. Automated outcome checking may offer a scalable way to support editors, peer reviewers, and authors in detecting outcome switching and improving trial reporting.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The use of the Registered Reports format for publication of randomized clinical trials: a cross-sectional study 95%
- Large language models for conducting systematic reviews: on the rise, but not yet ready for use – a scoping review 94%
- Results reporting for clinical trials led by medical universities and university hospitals in the Nordic countries was often missing or delayed 93%
Similar papers in this journal
- Evaluation of SURUS: a Named Entity Recognition System to Extract Knowledge from Interventional Study Records 94%
- Approaches in Analyzing Predictors of Trial Failure: A Scoping Review and Meta-epidemiological study 94%
- Improving research transparency with individualized report cards: A feasibility study in clinical trials at a large university medical center 94%
Similar papers in this journal
- A modular pipeline for natural language processing-screened human abstraction of a pragmatic trial outcome from electronic health records 93%
- Use of estimands in cluster randomised trials: a review 92%
- Dynamic methods for ongoing assessment of site-level risk in risk-based monitoring of clinical trials: a scoping review 92%
Similar papers in this journal
- INSIGHT: A Tool for Fit-for-Purpose Evaluation and Quality Assessment of Observational Data Sources for Real World Evidence on Medicine and Vaccine Safety 92%
- A Systematic Process for Assessing Fitness-for-Purpose of Health Outcomes for Computable Phenotyping with Electronic Health Record Data 92%
- Using Natural Language Processing of Clinical Notes to Supplement Structured Electronic Health Record Data for Phenotyping Smoking and Obesity in a Healthcare System 89%
Similar papers in this journal
- Evaluation of the sensitivity, accuracy and currency of the Cochrane COVID-19 Study Register for supporting rapid evidence synthesis production 95%
- Development of the Individual Participant Data (IPD) Integrity Tool for assessing the integrity of randomised trials using individual participant data 94%
- Catchii: empowering literature review screening in healthcare 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.