Performance of Heart Failure Clinical Prediction Models: A Systematic External Validation Study
Upshaw, J.; Nelson, J.; Koethe, B.; Park, J. G.; McGinnes, H. L.; Wessler, B. S.; Konstam, M. A.; Udelson, J. E.; Van Calster, B.; van Klaveren, D.; Steyerberg, E.; Kent, D. M.
Show abstract
BackgroundMost heart failure (HF) clinical prediction models (CPMs) have not been externally validated. MethodsWe performed a systematic review to identify CPMs predicting outcomes in HF, stratified by acute and chronic HF CPMs. External validations were performed using individual patient data from 8 large HF trials (1 acute, 7 chronic). CPM discrimination (c-statistic, % relative change in c-statistic), calibration (calibration slope, Harrells E, E90), and net benefit were evaluated for each CPM with and without recalibration. ResultsOf 135 HF CPMs screened, 24 (18%) were compatible with the population, predictors and outcomes to the trials and 42 external validations were performed (14 acute HF, 28 chronic HF). The median derivation c-statistic of acute HF CPMs was 0.76 (IQR, 0.75, 0.8), validation c-statistic was 0.67 (0.65, 0.68) and model-based c-statistic was 0.68 (0.66, 0.76), Hence, most of the apparent decrement in model performance was due to narrower case-mix in the validation cohort compared with the development cohort. The median derivation c-statistic for chronic HF CPMs was 0.76 (0.74, 0.8), validation c-statistic 0.61 (0.6, 0.63) and model-based c-statistic 0.68 (0.62, 0.71), suggesting that the decrement in model performance was only partially due to case-mix heterogeneity. Calibration was generally poor - median E (standardized by outcome rate) was 0.5 (0.4, 2.2) for acute HF CPMs and 0.5 (0.3, 0.7) for chronic HF CPMs. Updating the intercept alone led to a significant improvement in calibration in acute HF CPMs, but not in chronic HF CPMs. Net benefit analysis showed potential for harm in using CPMs when the decision threshold was not near the overall outcome rate but this improved with model recalibration. ConclusionsOnly a small minority of published CPMs contained variables and outcomes that were compatible with the clinical trial datasets. For acute HF CPMs, discrimination is largely preserved after adjusting for case-mix; however, the risk of net harm is substantial without model recalibration for both acute and chronic HF CPMs.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- External validation of a claims-based model to predict left ventricular ejection fraction class in patients with heart failure 97%
- Predicting 30-Day and 1-Year Mortality in Heart Failure with Preserved Ejection Fraction (HFpEF) 95%
- Hemodynamic profiles by non-invasive monitoring of cardiac index and vascular tone in acute heart failure patients in the emergency department: external validation and clinical outcomes 94%
Similar papers in this journal
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 94%
- Development and Validation of a parsimonious AI-Based Risk Score for Mortality in Heart Failure: A UK cohort study 94%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 94%
Similar papers in this journal
- Arrhythmia and Survival Outcomes among Black and White Patients with a Primary Prevention Defibrillator 95%
- Aptamer Proteomics for Biomarker Discovery in Heart Failure with Reduced Ejection Fraction 93%
- Screening for Atrial Fibrillation in Older Adults at Primary Care Visits: the VITAL-AF Randomized Controlled Trial 93%
Similar papers in this journal
- Multiple Biomarkers to Predict Major Adverse Cardiovascular Events in Patients With Coronary Chronic Total Occlusions 95%
- A Multicenter Evaluation of the Impact of Procedural and Pharmacological Interventions on Deep Learning-based Electrocardiographic Markers of Hypertrophic Cardiomyopathy 94%
- Risk of Cardiovascular Events after Covid-19: a double-cohort study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.