An Artificial Intelligence Model for Detection of Heart Failure with Preserved Ejection Fraction: A Report from HeartShare Study
Karabayir, I.; Singh, S.; Hayit, T.; Soliman, E. Z.; Kitzman, D.; Herrington, D. M.; Borlaug, B. A.; Davis, R. L.; Shah, S. J.; Akbilgic, O.
Show abstract
BackgroundHeart failure with preserved ejection fraction (HFpEF) accounts for over half of all heart failure cases in the United States and remains a diagnostic challenge. Non-invasive, scalable screening tools may enable earlier recognition, timely intervention, and improved care. To evaluate the performance, reproducibility, and early detection capability of an electrocardiogram-based artificial intelligence (ECG-AI) model designed to identify HFpEF using HeartShare data and real-world ECGs from Wake Forest Baptist Health (WFBH). MethodsThe original ECG-AI model was developed and validated using >1 million ECGs. In this study, we examined the external validity and reproducibility over time of this ECG-AI measure in 432 participants from an NIH-funded study of clinically validated HFpEF or controls (HeartShare). Specifically, we assessed model accuracy (AUC, sensitivity, specificity, predictive values) and reproducibility across three serial ECGs. We also analyzed the potential for early (preclinical) detection of HFpEF in 59,705 real-world ECGs from 12,338 patients a large integrated healthcare system (Wake Forest Baptist Health (WFBH)). ResultsIn HeartShare, ECG-AI achieved an AUC of 0.760 (95% CI: 0.729-0.816), with 65% sensitivity and 75% specificity for detection of HFpEF. We obtained no significantly different AUC when using only lead I ECG as an input, AUC of 0.773 (0.729-0.816). Within-patient reproducibility across three consecutive ECGs showed strong correlations (Pearson r = 0.87-0.89) and strong agreement (Cohens {kappa} = 0.68-0.74). Misclassified cases showed fewer risk factors and more normal-like ECG features. In real-world WFBH data, ECG-AI detected HFpEF up to 4 years before clinical diagnosis with AUCs from 0.77 to 0.80. Conclusions12 lead ECG-AI model demonstrates strong generalizability, reproducibility and early detection capabilities for HFpEF, supporting its potential as a scalable screening and risk stratification tool. Almost identical single lead AUC demands future investigation for remote monitoring. What is new?This study is the first to demonstrate that an ECG-AI model for HFpEF maintains strong temporal reproducibility across serial ECGs, supporting its stability as a robust non-invasive tool. We validate the model in both a rigorously phenotyped cohort and a large real-world health system and show that single-lead ECG input achieves accuracy comparable to the full 12-lead model. In addition, we show that the model can identify HFpEF years before clinical diagnosis, extending prior work by establishing ECG-AI as a reproducible, generalizable, and potentially preclinical detection tool. What are the clinical implications?The strong temporal reproducibility of the ECG-AI measure indicates that it can provide reliable longitudinal tracking of HFpEF risk, making it suitable for both clinical monitoring and remote assessment. Early detection capabilities - up to four years before diagnosis - create opportunities for proactive evaluation and earlier intervention. The comparable performance of single-lead ECGs also opens the door for scalable deployment through wearable or home-based devices, broadening access to HFpEF screening and enabling continuous risk surveillance outside traditional clinical environments.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 96%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 96%
- Advanced ECG heart age estimation applicable to both sinus and non-sinus rhythm associates with cardiovascular risk, cardiovascular morbidity, and survival 95%
Similar papers in this journal
- Development and validation of imaging-free myocardial fibrosis prediction models, association with outcomes, and sample size estimation for phase 3 trials 95%
- Relationship of mild to moderate impairment of left ventricular ejection fraction with fatal ventricular arrhythmic events in cardiac sarcoidosis 95%
- An International Longitudinal Natural History Study of Danon Disease Patients: Unique Cardiac Trajectories Identified Based on Sex and Heart Failure Outcomes 94%
Similar papers in this journal
- Multiple Biomarkers to Predict Major Adverse Cardiovascular Events in Patients With Coronary Chronic Total Occlusions 95%
- Prognostic value of compact myocardial thinning in patients with left ventricular non-compaction 95%
- A Multicenter Evaluation of the Impact of Procedural and Pharmacological Interventions on Deep Learning-based Electrocardiographic Markers of Hypertrophic Cardiomyopathy 95%
Similar papers in this journal
- External validation of a claims-based model to predict left ventricular ejection fraction class in patients with heart failure 95%
- Predicting 30-Day and 1-Year Mortality in Heart Failure with Preserved Ejection Fraction (HFpEF) 95%
- Predictors and outcomes of Cardiac Dyssynchrony among patients with heart failure attending Benjamin Mkapa Hospital in Dodoma, central Tanzania: A protocol of prospective-longitudinal study 94%
Similar papers in this journal
- Artificial Intelligence Methods to Detect Heart Failure with Preserved Ejection Fraction (AIM-HFpEF) within Electronic Health Records: An equitable disease prediction model 94%
- Right Ventricular Dysfunction: An Overlooked Predictor of Sudden Cardiac Death and Arrhythmic Events -- A Meta-Analysis 93%
- Clinical presentation, disease course and outcome of COVID-19 in hospitalized patients with and without pre-existing cardiac disease – a cohort study across sixteen countries 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.