Natural Language Processing for Adjudication of Heart Failure Hospitalizations in a Multi-Center Clinical Trial
Cunningham, J. W.; Singh, P.; Reeder, C.; Claggett, B.; Marti-Castellote, P. M.; Lau, E. S.; Khurshid, S.; Batra, P.; Lubitz, S.; Maddah, M.; Philippakis, A.; Desai, A.; Ellinor, P. T.; Vardeny, O.; Solomon, S. D.; Ho, J. E.
Show abstract
BackgroundThe gold standard for outcome adjudication in clinical trials is chart review by a physician clinical events committee (CEC), which requires substantial time and expertise. Automated adjudication by natural language processing (NLP) may offer a more resource-efficient alternative. We previously showed that the Community Care Cohort Project (C3PO) NLP model adjudicates heart failure (HF) hospitalizations accurately within one healthcare system. MethodsThis study externally validated the C3PO NLP model against CEC adjudication in the INVESTED trial. INVESTED compared influenza vaccination formulations in 5260 patients with cardiovascular disease at 157 North American sites. A central CEC adjudicated the cause of hospitalizations from medical records. We applied the C3PO NLP model to medical records from 4060 INVESTED hospitalizations and evaluated agreement between the NLP and final consensus CEC HF adjudications. We then fine-tuned the C3PO NLP model (C3PO+INVESTED) and trained a de novo model using half the INVESTED hospitalizations, and evaluated these models in the other half. NLP performance was benchmarked to CEC reviewer inter-rater reproducibility. Results1074 hospitalizations (26%) were adjudicated as HF by the CEC. There was high agreement between the C3PO NLP and CEC HF adjudications (agreement 87%, kappa statistic 0.69). C3PO NLP model sensitivity was 94% and specificity was 84%. The fine-tuned C3PO and de novo NLP models demonstrated agreement of 93% and kappa of 0.82 and 0.83, respectively. CEC reviewer inter-rater reproducibility was 94% (kappa 0.85). ConclusionOur NLP model developed within a single healthcare system accurately identified HF events relative to the gold-standard CEC in an external multi-center clinical trial. Fine-tuning the model improved agreement and approximated human reproducibility. NLP may improve the efficiency of future multi-center clinical trials by accurately identifying clinical events at scale.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Arrhythmia and Survival Outcomes among Black and White Patients with a Primary Prevention Defibrillator 94%
- Aptamer Proteomics for Biomarker Discovery in Heart Failure with Reduced Ejection Fraction 93%
- High-sensitivity cardiac troponin on presentation to rule out myocardial infarction: a stepped-wedge cluster randomised controlled trial 93%
Similar papers in this journal
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 96%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 94%
- International Evaluation Of An Artificial Intelligence-Powered Ecg Model Detecting Occlusion Myocardial Infarction 94%
Similar papers in this journal
- Comparative Effectiveness of Second-line Antihyperglycemic Agents for Cardiovascular Outcomes: A Large-scale, Multinational, Federated Analysis of the LEGEND-T2DM Study 95%
- ClinGen Hereditary Cardiovascular Disease Gene Curation Expert Panel: Reappraisal of Genes associated with Hypertrophic Cardiomyopathy 91%
- Lipid-Modulating Agents for Prevention or Treatment of COVID-19 in Randomized Trials 91%
Similar papers in this journal
- A Multicenter Evaluation of the Impact of Procedural and Pharmacological Interventions on Deep Learning-based Electrocardiographic Markers of Hypertrophic Cardiomyopathy 95%
- Natural Language Processing for the Ascertainment and Phenotyping of Left Ventricular Hypertrophy and Hypertrophic Cardiomyopathy on Echocardiogram Reports 94%
- Multiple Biomarkers to Predict Major Adverse Cardiovascular Events in Patients With Coronary Chronic Total Occlusions 94%
Similar papers in this journal
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 95%
- Identification of Digital Twins to Guide Interpretable AI for Diagnosis and Prognosis in Heart Failure 92%
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.