An ECG foundation model for generalizable cardiac function prediction across the lifespan
Yang, Y.; Peracchio, L.; Mayourian, J.; Miller, T.; La Cava, W.
Show abstract
Background Artificial intelligence-enhanced electrocardiography (AI-ECG) enables scalable, low-cost cardiac dysfunction screening, but existing models are annotation-intensive and predominantly adult-derived, leaving paediatric generalizability uncertain. Paediatric cohorts exhibit highly variable cardiac morphology and function compared to adults, which may be useful for learning generalizable AI-ECG models. Methods We pretrained ECG-Fyler on a predominantly paediatric, all-age cohort at Boston Children's Hospital (1992-2023), annotated with a cardiology-specific coding system (Fyler codes), and evaluated it on assessments from echocardiography (echo) and cardiac magnetic resonance (CMR) studies. We validated on an external adult cohort from Columbia University Irving Medical Center. Performance was benchmarked against several AI-ECG foundation models by AUROC across age groups, lesion types, and limited-data scenarios. Findings The pretraining cohort comprised 782,138 ECGs from 255,271 patients (median age: 10.9 years, IQR: [2.8-16.8]). Internal evaluation included 178,495 ECG-echo pairs (median age: 10.9 [3.7-17.0]) and 8,584 ECG-CMR pairs (median age: 20.7 [15.6-29.6]). External validation included 82,543 ECG-echo pairs from adults (median age: 64.0 [52.0-74.0]). ECG-Fyler improved AUROC across biventricular dysfunction and dilation tasks, with the largest gains in low-data settings. In internal validation, ECG-Fyler detected low left ventricular ejection fraction (LVEF [≤] 40%) from only 100 fine-tuning samples (AUROC: 0.80, 95% CI: [0.78-0.80]), outperforming other models (AUROC < 0.65) and improving with additional fine-tuning (AUROC: 0.94 [0.93-0.94]). Similar improvements were observed for CMR-derived LVEF, RVEF, and ventricular dilation. In external validation on adults, ECG-Fyler exhibited an AUROC of 0.83 (CI: [0.82-0.85]) for LVEF [≤] 40%. After fine-tuning on less than 10% of external data, LVEF [≤] 45% performance (AUROC: 0.87 [0.86-0.88]) outperformed a fully trained, site-specific prior model (AUROC: 0.85 [0.84-0.87]). Interpretation Pretraining on richly annotated, paediatric-dominant ECGs yields models that transfer efficiently across institutions and ages, supporting AI-ECG screening and triage when labels or imaging access are limited. Funding National Institutes of Health (R01LM012973); Kostin Innovation Fund, Boston Children's Hospital
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Multimodal deep learning enhances diagnostic precision in left ventricular hypertrophy 96%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 95%
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 95%
Similar papers in this journal
- Expert-level prenatal detection of complex congenital heart disease from screening ultrasound using deep learning 91%
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 90%
- Fostering transparent medical image AI via an image-text foundation model grounded in medical literature 88%
Similar papers in this journal
- Detecting QT prolongation From a Single-lead ECG With Deep Learning 92%
- Multiple Instance Learning Framework can Facilitate Explainability in Murmur Detection 91%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 91%
Similar papers in this journal
- Automated severe aortic stenosis detection on single-view echocardiography: A multi-center deep learning study 96%
- AORTA Gene: Polygenic prediction improves detection of thoracic aortic aneurysm 90%
- Deciphering human heart failure with preserved ejection fraction (HFpEF) at single cell resolution 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.