European Heart Journal - Digital Health
◐ Oxford University Press (OUP)
All preprints, ranked by how well they match European Heart Journal - Digital Health's content profile, based on 18 papers previously published here. The average preprint has a 0.04% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Aminorroaya, A.; Vasisht Shankar, S.; Carter, M.; Khan, M.; Dhingra, L. S.; Khunte, A.; Croon, P. M.; Lombo, B.; McNamara, R. L.; Oikonomou, E. K.; Pedroso, A. F.; Khera, R.
Show abstract
Importance: Consumer wearables such as the Apple Watch can record single-lead electrocardiograms (ECGs) but are used mainly to detect rhythm disorders. Artificial intelligence-enhanced ECG (AI-ECG) could extend these real-world recordings for detecting structural heart disease (SHD), yet prospective validation remains limited. Objective: To prospectively validate a previously developed, noise-adapted AI-ECG model for detecting severe SHD from single-lead Apple Watch ECGs. Design: Prospective cohort study. Setting: Yale New Haven Hospital echocardiography laboratory. Participants: Adults aged >=18 years undergoing outpatient transthoracic echocardiography (TTE) as part of routine clinical care. Exposure: A 30-second, single-lead Apple Watch ECG recorded during the TTE visit and processed through an end-to-end, HIPAA-compliant platform for real-time AI-ECG inference. Main Outcomes and Measures: The primary outcome was discrimination for TTE-defined severe SHD, a composite of left ventricular systolic dysfunction (left ventricular ejection fraction <40%), severe left-sided valvular disease, and/or severe left ventricular hypertrophy, assessed by the area under the receiver operating characteristic curve (AUROC). Secondary measures were sensitivity, specificity, negative predictive value (NPV), and positive predictive value (PPV) at prespecified thresholds, and screening efficiency, assessed by the number needed to test (NNT) under usual-care versus AI-ECG-guided strategies. Results: Among 596 participants with analyzable Apple Watch ECGs (median age, 62 years [IQR, 46-72]; 51.2% women), severe SHD was present in 30 (5.1%). The model discriminated severe SHD well (AUROC, 0.841; 95% CI, 0.761-0.921), with a sensitivity of 76.7% (59.1-88.2), specificity of 83.2% (79.9-86.1), NPV of 98.5% (97.0-99.3), and PPV of 19.7% (13.5-27.8) at the prespecified threshold. An AI-ECG-guided strategy reduced the NNT to identify one case by more than 60% versus usual care across the composite and individual SHD phenotypes. Conclusions and Relevance: In this prospective cohort, a noise-adapted AI-ECG algorithm identified SHD phenotypes from real-world single-lead Apple Watch ECGs and improved screening efficiency. These findings support a potential role for wearable ECG-based screening in the scalable identification of clinically actionable SHD.
Knight, E.; Oikonomou, E. K.; Aminorroaya, A.; Pedroso, A. F.; Khera, R.
Show abstract
Artificial intelligence (AI) models can now detect patterns of structural heart diseases (SHDs) from electrocardiograms (ECGs), though scaling them requires the broader use of single-lead ECGs that are now ubiquitous in wearable and portable devices. However, model development for these devices is limited by a lack of diagnostic labels for SHDs for wearable ECGs. Here, we present Wearable-Echo-FM, a foundation model that encodes single-lead ECGs with information from echocardiographic text reports. Using 274,057 single-lead ECG-echo pairs from 77,378 adults (2015-2019), we contrastively pre-trained convolutional neural network (CNN) and RoBERTa encoders. The ECG encoder was fine-tuned on a distinct progressively larger ECG set (250 to 250,260 ECGs) to detect different cardiac disorders (i) left-ventricular systolic dysfunction (LVSD), (ii) diastolic dysfunction, and (iii) a composite SHD. This was compared with a randomly initialized CNN, with both approaches evaluated in an independent held-out test set. With the full training set, Wearable-Echo-FM matched the baseline CNN (AUROC 0.894 vs 0.884 for LVSD; 0.849 vs 0.843 diastolic dysfunction; 0.887 vs 0.869 composite). With only 0.5% (~1000 ECGs) of data, it markedly outperformed baseline (0.855 vs 0.548; 0.819 vs 0.582; 0.863 vs 0.496, respectively). Contrastive pre-training of single-lead ECGs on echocardiographic text reduces label requirements for SHD screening on wearable and portable devices.
Dhingra, L. S.; Aminorroaya, A.; Sangha, V.; Pedroso Camargos, A.; Vasisht Shankar, S.; Coppi, A.; Foppa, M.; Brant, L. C. C.; Barreto, S. M.; Ribeiro, A. L. P.; Krumholz, H.; Oikonomou, E. K.; Khera, R.
Show abstract
BackgroundIdentifying structural heart diseases (SHDs) early can change the course of the disease, but their diagnosis requires cardiac imaging, which is limited in accessibility. ObjectiveTo leverage images of 12-lead ECGs for automated detection and prediction of multiple SHDs using an ensemble deep learning approach. MethodsWe developed a series of convolutional neural network models for detecting a range of individual SHDs from images of ECGs with SHDs defined by transthoracic echocardiograms (TTEs) performed within 30 days of the ECG at the Yale New Haven Hospital (YNHH). SHDs were defined as LV ejection fraction <40%, moderate-to-severe left-sided valvular disease (aortic/mitral stenosis or regurgitation), or severe left ventricular hypertrophy (IVSd > 1.5cm and diastolic dysfunction). We developed an ensemble XGBoost model, PRESENT-SHD, as a composite screen across all SHDs. We validated PRESENT-SHD at 4 US hospitals and the prospective, population-based Brazilian Longitudinal Study of Adult Health (ELSA-Brasil), with concurrent protocolized ECGs and TTEs. We also used PRESENT-SHD for risk stratification of new-onset SHD or heart failure (HF) in clinical cohorts and the population-based UK Biobank (UKB). ResultsThe models were developed using 261,228 ECGs from 93,693 YNHH patients and evaluated on a single ECG from 11,023 individuals at YNHH (19% with SHD), 44,591 across external hospitals (20-27% with SHD), and 3,014 in the ELSA-Brasil (3% with SHD). In the held-out test set, PRESENT-SHD demonstrated an AUROC of 0.886 (0.877-894), 90% sensitivity, and 66% specificity. At hospital-based sites, PRESENT-SHD had AUROCs ranging from 0.854-0.900, with sensitivities and specificities of 93-96% and 51-56%, respectively. The model generalized well to ELSA-Brasil (AUROC, 0.853 [0.811-0.897], 88% sensitivity, 62% specificity). PRESENT-SHD demonstrated consistent performance across demographic subgroups, novel ECG formats, and smartphone photographs of ECGs from monitors and printouts. A positive PRESENT-SHD screen portended a 2- to 4-fold higher risk of new-onset SHD/HF, independent of demographics, comorbidities, and the competing risk of death across clinical sites and UKB, with high predictive discrimination. ConclusionWe developed and validated PRESENT-SHD, an AI-ECG tool identifying a range of SHD using images of 12-lead ECGs, representing a robust, scalable, and accessible modality for automated SHD screening and risk stratification. CONDENSED ABSTRACTScreening for structural heart disorders (SHDs) requires cardiac imaging, which has limited accessibility. To leverage 12-lead ECG images for automated detection and prediction of multiple SHDs, we developed PRESENT-SHD, an ensemble deep learning model. PRESENT-SHD demonstrated excellent performance in detecting SHDs across 5 US hospitals and a population-based cohort in Brazil. The model successfully predicted the risk of new-onset SHD or heart failure in both US clinical cohorts and the community-based UK Biobank. By using ubiquitous ECG images and smartphone photographs to predict a composite outcome of multiple SHDs, PRESENT-SHD establishes a scalable paradigm for cardiovascular screening and risk stratification.
Dhingra, L. S.; Croon, P. M.; Batinica, B.; Aminorroaya, A.; Pedroso, A. F.; Oikonomou, E. K.; Khera, R.
Show abstract
BackgroundThe scientific literature on artificial intelligence-enabled electrocardiography (AI-ECG) has defined a robust performance of AI models in detecting and predicting several structural heart disorders (SHDs) using ECGs. However, as a diagnostic test, the real-world clinical utility of AI-ECG reliability requires the consistency of its results when repeated under similar conditions. AimTo evaluate the reliability of AI-ECG models for different ECGs for the same person, across different diagnostic labels, and using varied modeling approaches. MethodsWe used ECG images (2000-2024) from 5 hospitals and an outpatient network within a large, integrated US health system. For each individual, we identified multiple ECGs recorded within a 30-day period. We evaluated 7 models: 6 convolutional neural networks (CNNs) trained to detect individual SHDs, including LV systolic dysfunction, left valve diseases and severe LVH; an ensemble XGBoost integrating individual CNNs as a composite screen for multiple SHDs. We used concordance correlation coefficient (CCC), Spearman correlation, Cohens kappa, and percent agreement in binary screen status to test model reliability. We evaluated factors associated with different AI-ECG outputs ({Delta} probability> 0.5) and assessed stability across ECG layouts (digital, printed, photo). ResultsAcross sites, we identified 1,118,263 ECG pairs, with a median 1 (1-3) days between ECGs. The ensemble XGBoost had the higher test-retest correlation (CCC: 0.89-0.92) and agreement (kappa: 0.75-0.82) between pairs compared with CNNs (CCC: 0.78-0.88; kappa: 0.57-0.72). After adjusting for demographics, ECG pairs that included one or both inpatient ECG were significantly more likely to yield unstable predictions (ORs: 1.60 [1.50-1.70] and 1.91 [1.78-2.05], respectively) compared with pairs with both ECGs obtained in outpatient settings. Among outpatient pairs across sites, the XGBoost model had a CCC of 0.89-0.94, a Spearman correlation of 0.90-0.94, and a kappa of 0.78-0.84, with concordance rates of 89-92%. Notably, ensemble model predictions were also stable across different ECG layouts. ConclusionAn ensemble AI-ECG model integrating multiple CNN predictions had higher reliability compared with models for individual disorders. Discordance was more common in inpatient ECGs, suggesting instability in high-acuity settings. Reliable ensemble AI-ECG model outputs support readiness for clinical implementation for SHD screening. GRAPHICAL ABSTRACTO_ST_ABSStudy DesignC_ST_ABSAbbreviations: AR, aortic regurgitation; AS, aortic stenosis; CNN, convolutional neural network; ECG, electrocardiogram; FC, fully-connected layers; LVSD, left ventricular systolic dysfunction; MR, mitral regurgitation; SHD, structural heart diseases; sLVH, severe left ventricular hypertrophy, XGBoost, extreme gradient boosting. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=140 SRC="FIGDIR/small/25339526v1_ufig1.gif" ALT="Figure 1"> View larger version (45K): org.highwire.dtl.DTLVardef@7aaba3org.highwire.dtl.DTLVardef@19a8bd1org.highwire.dtl.DTLVardef@151516borg.highwire.dtl.DTLVardef@1b8738f_HPS_FORMAT_FIGEXP M_FIG C_FIG
van de Leur, R. R.; Bos, M. N.; Taha, K.; Sammani, A.; van Duijvenboden, S.; Lambiase, P.; Hassink, R. J.; van der Harst, P.; Doevendans, P. A.; Gupta, D.; van Es, R.
Show abstract
BackgroundDeep neural networks (DNNs) show excellent performance in interpreting electrocardiograms (ECGs), both for conventional ECG interpretation and for novel applications such as detection of reduced ejection fraction and prediction of one-year mortality. Despite these promising developments, clinical implementation is severely hampered by the lack of trustworthy techniques to explain the decisions of the algorithm to clinicians. Especially, currently employed heatmap-based methods have shown to be inaccurate. MethodsWe present a novel approach that is inherently explainable and uses an unsupervised variational auto-encoder (VAE) to learn the underlying factors of variation of the ECG (the FactorECG) in a database with 1.1 million ECG recordings. These factors are subsequently used in a pipeline with common and interpretable statistical methods. As the ECG factors are explainable by generating and visualizing ECGs on both the model- and individual patient-level, the pipeline becomes fully explainable. The performance of the pipeline is compared to a state-of-the-art black box DNN in three tasks: conventional ECG interpretation with 35 diagnostic statements, detection of reduced ejection fraction and prediction of one-year mortality. FindingsThe VAE was able to compress the ECG into 21 generative ECG factors, which are associated with physiologically valid underlying anatomical and (patho)physiological processes. When applying the novel pipeline to the three tasks, the explainable FactorECG pipeline performed similar to state-of-the-art black box DNNs in conventional ECG interpretation (AUROC 0{middle dot}94 vs 0{middle dot}96), detection of reduced ejection fraction (AUROC 0{middle dot}90 vs 0{middle dot}91) and prediction of one-year mortality (AUROC 0{middle dot}76 vs 0{middle dot}75). Contrary to state-of-the-art, our pipeline provided inherent explainability on which morphological ECG features were important for prediction or diagnosis. InterpretationFuture studies should employ DNNs that are inherently explainable to facilitate clinical implementation by gaining confidence in artificial intelligence, and more importantly, making it possible to identify biased or inaccurate models. FundingThis study was financed by the Netherlands Organisation for Health Research and Development (ZonMw, no. 104021004) and the Dutch Heart Foundation (no. 2019B011). Research into ContextO_ST_ABSEvidence before this studyC_ST_ABSA comprehensive literature survey was performed for research articles on interpretable or explainable artificial intelligence (AI) for interpretation of raw electrocardiograms (ECGs) using PubMed and Google Scholar databases. Articles in English up to November 24, 2021, were included and the following key words were used: deep neural network (DNN), deep learning, convolutional neural network, artificial intelligence, electrocardiogram, ECG, explainability, explainable, interpretability, interpretable, and visualization. Many studies that used DNNs to interpret ECGs with high predictive performances were found, some focusing on tasks known to be associated with the ECG (e.g., rhythm disorders) and others identifying completely novel use cases for the ECG (e.g. reduced ejection fraction). All of these studies employed post-hoc explainability techniques, where the decisions of the black box DNN were visualized after training, usually using heatmaps (i.e., using Grad-CAM, SHAP or LIME). In these studies, only some example ECGs were handpicked, as these heatmap-based techniques only work on single ECGs. Three studies also investigated the global features of the model by taking a summary measure of the heatmaps, by relating heatmaps to known ECG parameters (i.e., QRS duration) or by using prototypes. No studies investigated whether the features found using heatmaps were robust or reproducible. Added value of this studyCurrently employed post-hoc explainability techniques, usually heatmap-based, have limited explainable value as they merely indicate the temporal location of a specific feature in the individual ECG. Moreover, these techniques have been shown to be unreliable, poorly reproducible and suffer from confirmation bias. To address this gap in knowledge, we designed a DNN that is inherently explainable (i.e. explainable by design instead of investigating post-hoc). This DNN is used in a pipeline that consists of three components: (i) a generative DNN (variational auto-encoder) that learned to encode the ECG into its underlying 21 continuous factors of variation (the FactorECG), (ii) a visualization technique to provide insight into these ECG factors, and (iii) a common interpretable statistical method to perform diagnosis or prediction using the ECG factors. Model-level explainability is obtained by varying the ECG factors while generating and plotting ECGs, which allows for visualization of detailed changes in morphology, that are associated with physiologically valid underlying anatomical and (patho)physiological processes. Moreover, individual patient-level explanations are also possible, as every individual ECG has its representative set of explainable FactorECG values, of which the associations with the outcome are known. When using the explainable pipeline for interpretation of diagnostic ECG statements, detection of reduced ejection fraction and prediction of one-year mortality, it yielded predictive performances similar to state-of-the-art black box DNNs. Contrary to the state-of-the-art, our pipeline provided inherent explainability on which ECG features were important for prediction or diagnosis. For example, ST elevation was discovered to be an important predictor for reduced ejection fraction, which is an important finding as it could limit the generalizability of the algorithm to the general population. Implications of all the available evidenceA longstanding assumption was that the high-dimensional and non-linear black box nature of the currently applied ECG-based DNNs was inevitable to gain the impressive performances shown by these algorithms on conventional and novel use cases. This study, however, shows that inherently explainable DNNs should be the future of ECG interpretation, as they allow reliable clinical interpretation of these models without performance reduction, while also broadening their applicability to detect novel features in many other (rare) diseases. The application of such methods will lead to more confidence in DNN-based ECG analysis, which will facilitate the clinical implementation of DNNs in routine clinical practice.
Meneguitti Dias, F.; Ribeiro, E.; Olivetti, N.; Carvalho, O.; Krieger, J. E.; Gutierrez, M.
Show abstract
Automated electrocardiogram analysis has advanced largely through digital waveforms, yet many emergency-care workflows rely on ECGs available only as printed tracings, scanned reports, PDFs or mobile photographs. We developed an image-based deep learning system for emergency ECG classification and evaluated it in InCor-EMG, an expert-adjudicated dataset of 18,519 emergency ECGs spanning 12 ECG categories, with labels from 19 cardiologists. On the held-out test set, the final ConvNeXt ensemble achieved a macro F1-score of 0.807 (95% CI, 0.788-0.825), compared with 0.820 (95% CI, 0.805-0.832) for annotating cardiologists, and higher F1-scores than Mortara Veritas in most evaluated categories. Performance was associated more strongly with inter-reader agreement than with training sample size and remained informative across scanned and photographed ECGs, with supportive performance in model-enriched temporal and heterogeneous public-image evaluations. These findings support ECG image classification when digital waveforms are unavailable.
Pitre, T.; Marques, L.; Weatherald, J.; Mak, S.; Thavendiranathan, P.; Granton, J.
Show abstract
Background: Right ventricular (RV) function predicts survival in pulmonary hypertension (PH) and other cardiovascular diseases, yet echocardiographic AI has largely focused on the left ventricle (LV). Objectives: To develop and evaluate PH-ECHO-AI, a unified deep learning model performing four-chamber segmentation, landmark localisation, biventricular ejection fraction (EF) estimation, deformation analysis, and PH prediction from a single apical four-chamber (A4C) clip. Methods: We developed the model using 8,416 clips from four public datasets and no institutional data: EchoNet-Dynamic, CAMUS, RVENet (apical four-chamber clips paired with 3D-echocardiographic right ventricular ejection fraction, RVEF), and MIMIC-IV-ECHO. Evaluation used held-out, training-excluded data with expert-reviewed reference standards and a per-cohort audit of patient-level separation: 1,416 clips for segmentation; 600 clips for function and deformation (350 referenced to 3D-echocardiographic RVEF, 250 to the EchoNet LVEF); and 1,076 MIMIC-IV patients for PH prediction, with five-fold cross-validation. Performance measures were Dice, correlation, mean absolute error (MAE), Bland-Altman agreement, and area under the receiver operating characteristic curve (AUC). Results: Four-chamber segmentation generalised robustly across all datasets (pooled Dice: LV 0.925, RV 0.836, LA 0.910, RA 0.904). Left ventricular ejection fraction (LVEF) was estimated with r=0.845 (95% CI 0.786 to 0.886) and MAE 4.67%. RVEF, regressed directly from the clip by a supervised head trained on 3D-echocardiographic labels with no geometric assumption, reached r=0.754 (95% CI 0.690 to 0.806) and MAE 4.98%, matching published single-view RVEF ceilings and exceeding geometric RV fractional area change (RVFAC; r=0.278). Deformation and excursion metrics, namely RV free-wall and LV A4C longitudinal strain and tricuspid and mitral annular plane systolic excursion (TAPSE, MAPSE), proved physiologically coherent. Segmentation generalised to the external MIMIC-IV cohort, and PH prediction was developed and evaluated entirely within it; RVEF evaluation was clip-disjoint and same-source, so cross-centre RVEF validation remains outstanding. Using echocardiographic geometry alone, confirmed PH was detected with an AUC of 0.697 and strong calibration (Brier 0.061). Conclusions: A single, reproducible model provides comprehensive right-heart-focused interpretation from one A4C view. It achieves RVEF accuracy competitive with dedicated RV models while simultaneously delivering segmentation, deformation, annular excursion (TAPSE and MAPSE), and PH prediction. Registration: This retrospective study used existing datasets. Code is openly released, and trained model weights are available to credentialed investigators, for independent evaluation.
Croon, P. M.; Boonstra, M. J.; Allaart, C. P.; Arends, B. K. O.; Dhingra, L. S.; Huang, Y.-C.; Mast, T.; Khera, R.; Kuo, C.-F.; Kwon, J.-M.; Lee, H. S.; Lee, M. S.; van de Leur, R.; Liu, Z.-Y.; Oikonomou, E. K.; Selder, J. L.; Winter, M. M.; Asselbergs, F. W.
Show abstract
BackgroundSeveral artificial intelligence-enhanced electrocardiogram (AI-ECG) models have shown promise in detecting left ventricular systolic dysfunction (LVSD), but their head-to-head agreement and performance have not been independently compared within the same cohort. ObjectivesTo compare the performance of published AI-ECG models for LVSD detection in a standardized external cohort and evaluate the fields transparency and reproducibility. MethodsWe systematically reviewed AI-ECG models predicting LVSD and assessed the risk of bias. Authors were invited to share models for external validation in a well-phenotyped registry of patients undergoing routine clinical cardiac magnetic resonance imaging (CMR) with cardiologist-adjudicated reports and paired ECGs. Model performance was evaluated in all consecutive patients and a lower-complexity subgroup with 15% LVSD prevalence. ResultsWe identified 35 studies describing 51 models, reporting high (AUROC >0.80) or excellent (AUROC >0.90) performance. The risk of bias is high and primarily attributed to the limited description of development and validation cohort characteristics, as well as the lack of independent external validation. Four groups (from Korea, the United States, Taiwan, and the Netherlands) shared models for independent testing. AUROCs ranged from 0.83 to 0.93 in all patients (n = 1,203; mean age 59 {+/-} 15 years; 450 [35%] female) and from 0.87 to 0.96 in the lower complexity subset. Performance remained consistent across subgroups, with slight decreases in ECGs showing wide QRS complexes or atrial fibrillation. ConclusionsIn this first-in-kind independent validation and head-to-head comparison study, AI-ECG for LVSD detection demonstrated strong performance despite training on disparate populations. However, the limited availability of models hinders independent validation.
Aminorroaya, A.; Dhingra, L. S.; Sangha, V.; Oikonomou, E. K.; Khunte, A.; Shankar, S. V.; Camargos, A. P.; Haynes, N.; Hofer, I.; Ouyang, D.; Nadkarni, G.; Khera, R.
Show abstract
BackgroundDue to the lack of a feasible screening strategy, aortic stenosis (AS) is often diagnosed after the development of clinical symptoms, representing advanced stages of disease. Portable and wearable devices capable of recording electrocardiograms (ECGs) can be used for scalable screening for AS, if the diagnosis can be made with a single-lead ECG, despite potentially noisy acquisition. MethodsUsing electronic health records and imaging data from a large, diverse hospital system (2015-2022), we developed a deep learning-based approach to detect moderate/severe AS using a single-lead ECG. We used ECGs paired with echocardiograms obtained within 30 days of each other to develop the model. We extracted lead I signal data from clinical ECG and augmented it with random Gaussian noise. We trained a convolutional neural network (CNN) to identify TTE-confirmed AS using noisy single-lead ECGs. Finally, we used the CNN model probabilities, along with patient age and sex, as predictive inputs to train an extreme gradient boosting (XGBoost) model to detect moderate/severe AS. ResultsThe model was developed in 75,901 ECGs/35,992 patients (median age 61 [interquartile range (IQR) 47-72] years, 54.3% women, 9.5% Black) and validated in 3,733 patients (median age 61 [IQR 47-72] years, 53.4% women, 9.7% Black). In the held-out validation set, the ensemble XGBoost model achieved an AUROC of 0.829 (95% CI: 0.800-0.855), with a sensitivity of 90.4% and specificity of 58.7% for detecting moderate/severe AS. For detecting severe AS, the models AUROC was 0.846 (95% CI, 0.778-0.899), with a sensitivity of 94.3% and specificity of 57.0%. In the test set with a 4.5% prevalence of moderate/severe AS, the model had a PPV of 9.3% and an NPV of 99.2%. In simulated cohorts with 1% and 20% prevalence of moderate/severe AS, the models NPVs varied from 99.8% to 96.1%, and PPV from 2.2% to 35.4%, respectively. ConclusionWe developed a novel portable- and wearable-adapted deep learning approach for the detection of moderate/severe AS from noisy single-lead ECGs. Our approach represents a highly sensitive, feasible, and scalable strategy for community-based AS screening.
Khunte, A.; Sangha, V.; Oikonomou, E. K.; Dhingra, L. S.; Aminorroaya, A.; Coppi, A.; Vasisht Shankar, S.; Mortazavi, B. J.; Bhatt, D. L.; Krumholz, H. M.; Nadkarni, G.; Vaid, A.; Khera, R.
Show abstract
BackgroundTimely and accurate assessment of electrocardiograms (ECGs) is crucial for diagnosing, triaging, and clinically managing patients. Current workflows rely on computerized ECG interpretation tools built into ECG signal acquisition systems, which use rule-based algorithms that are unreliable and frequently not available in low-resource settings. We developed and validated a format-independent vision encoder-decoder model - ECG-GPT - that can generate free-text, expert-level interpretations directly from 12-lead ECG images. MethodsUsing 12-lead ECGs and their corresponding diagnosis statements collected at the Yale-New Haven Health System (YNHHS) between 2000 and 2022, we developed a vision-text transformer model to generate interpretation statements from images of ECGs. Using structured clinical assessment, semantic similarity, and conventional natural language generation metrics, we validated ECG-GPT across 7 geographically distinct health settings. These include (1) 3 large and diverse US health systems, (2) consecutive ECGs from a central reading system in Minas Gerais, Brazil, (3) the prospective cohort study, UK Biobank, (4) a Germany-based, publicly available repository, PTB-XL, and (5) a community hospital in Missouri. ResultsOverall, 2.9 million ECGs were used for model development. The model performed well in clinical assessment across 26 extracted labels: for atrial fibrillation, sinus tachycardia, sinus bradycardia, premature atrial contractions, and premature ventricular contractions, AUROCs and AUPRCs ranged from 0.80-0.95 and 0.50-0.86, respectively. For left bundle branch block, right bundle branch block, first degree atrioventricular block, left anterior fascicular block, and left posterior fascicular block, AUROCs and AUPRCs ranged from 0.88-0.96 and 0.23-0.86, respectively. Across all 26 conditions, diagnostic accuracy ranged between 0.93-0.99. ECG-GPT identified the full context of the diagnosis statements with allied conditions. It had a median pairwise cosine similarity of 0.90 (IQR 0.83-0.97), significantly greater than the median baseline similarity of 0.73 (IQR 0.67-0.78, p<0.001). This separation between median pairwise and baseline similarity remained consistent across all 26 condition-specific subsets. The results were comparable across external validation sites. ConclusionsWe developed and extensively validated a vision encoder-decoder model that generates expert-level interpretations from ECG images. This represents a scalable and accessible strategy for automated ECG analysis, especially in low-resource settings. CLINICAL PERSPECTIVEO_ST_ABSWhat is New?C_ST_ABSO_LIECG-GPT is a vision encoder-decoder model capable of generating full-text ECG interpretations directly from ECG images, regardless of layout or format. C_LIO_LIThe model was trained on over 2.7 million ECGs and externally validated across 3.8 million additional ECGs from demographically and geographically diverse populations. C_LI What are the clinical implications?O_LIECG-GPT enables automated, expert-level ECG interpretation directly from images, eliminating the need for signal data or device integration. C_LIO_LIThe model demonstrates consistent performance across diverse patient populations, ECG formats, and care settings. C_LIO_LIThis scalable, image-based approach may expand access to accurate ECG interpretation in low-resource settings. C_LI
Schlesinger, D.; Alam, R.; Ringel, R.; Pomerantsev, E.; Devreddy, S.; Shah, P.; Garasic, J.; Stultz, C.
Show abstract
BackgroundThe ability to non-invasively measure left atrial pressure would facilitate the identification of patients at risk of pulmonary congestion and guide proactive heart failure care. Wearable cardiac monitors, which record single-lead electrocardiogram data, provide information that can be leveraged to infer left atrial pressures. MethodsWe developed a deep neural network using single-lead electrocardiogram data to determine when the left atrial pressure is elevated. The model was developed and internally evaluated using a cohort of 6739 samples from the Massachusetts General Hospital (MGH) and externally validated on a cohort of 4620 samples from a second institution. We then evaluated model on patch-monitor electrocardiographic data on a small prospective cohort. ResultsThe model achieves an area under the receiver operating characteristic curve of 0.80 for detecting elevated left atrial pressures on an internal holdout dataset from MGH and 0.76 on an external validation set from a second institution. A further prospective dataset was obtained using single-lead electrocardiogram data with a patch-monitor from patients who underwent right heart catheterization at MGH. Evaluation of the model on this dataset yielded an area under the receiver operating characteristic curve of 0.875 for identifying elevated left atrial pressures for electrocardiogram signals acquired close to the time of the right heart catheterization procedure. ConclusionsThese results demonstrate the utility and the potential of ambulatory cardiac hemodynamic monitoring with electrocardiogram patch-monitors. Plain Language SummaryHeart failure is a prevalent disorder that is challenging to manage. Part of what makes heart failure management challenging is that the onset of symptoms can be insidious as there are few robust tools that a clinical can leverage to estimate when a patient is likely to experience an episode of acute heart failure. While elevated intracardiac pressures are a reliable harbinger of worsening heart failure, these pressures are typically measured using an invasive, gold-standard approach, which can only be performed in an inpatient setting. A non-invasive method for detecting higher pressures inside of heart would be especially helpful for identifying worsening heart failure in an expeditious manner in the home environment. For this reason, we developed an artificial intelligence model to detect elevated pressures inside the heart using a non-invasive signal, the electrocardiogram (ECG, or EKG), which can be acquired from a wearable patch monitor device. Our results demonstrate that the model provides a reliable platform for the non-invasive assessment of cardiac pressures using data that can be obtained in the outpatient setting.
Nicolson, A.; Pröll, S.; Lunelli, R.; Blankenburg, H.; Pramstaller, P.; Fuchsberger, C.; Bauer, A.; Dlaska, C.
Show abstract
Background Recent artificial intelligence (AI) models applied to the electrocardiogram (ECG) for risk stratification typically rely on supervised learning, defining risk as the error relative to an external target such as age or sex. This couples the risk score to the choice of target rather than the cardiac signal alone, and may limit generalisability. We aimed to develop a self-supervised AI-ECG risk score based on the error in reconstructing a partially masked ECG. Methods A transformer-based masked autoencoder was trained on 85% of the CODE dataset (n = 7,212,109 ECGs) to reconstruct ECG signals from partially masked inputs. The association between reconstruction error and all-cause mortality was assessed internally in CODE-15% and externally validated in four independent cohorts: MIMIC-IV-ECG (critical care, US), HEEDB (hospital, US), CHRIS (population-based, Italy), and Innsbruck (cardiology centre, Austria). A binary risk score (>1 SD above the CODE-15% mean) was additionally evaluated in these cohorts and in the UK Biobank (population-based, UK). Findings In Cox proportional hazards models adjusted for age and sex, each 1-SD increase in reconstruction error was associated with higher all-cause mortality (all p<0.001; cohort median follow-up 1.4-11.0 years): CODE-15% (HR 1.39, 95% CI 1.37-1.42), MIMIC-IV-ECG (HR 1.39, 95% CI 1.37-1.40), HEEDB (HR 1.41, 95% CI 1.40-1.41), Innsbruck (HR 1.23, 95% CI 1.21-1.26), and CHRIS (HR 1.25, 95% CI 1.14-1.38). The binary threshold identified a high-risk group with increased mortality in all six cohorts, including the UK Biobank (HR 1.27, 95% CI 1.08-1.50, p=0.004). Interpretation Reconstruction error is a generalisable predictor of all-cause mortality across diverse clinical and population-based settings. Unlike supervised approaches, it reflects the model's uncertainty about the ECG signal itself rather than error relative to an external target, providing a direct measure of how much each recording deviates from normal cardiac electrical patterns.
Aydogdu, D.; Gaber, F.; Sorooshmehr, A.; Akalin, A.
Show abstract
Cardiovascular diseases (CVDs) remain the primary global health burden, motivating the search for robust, non-invasive risk biomarkers. We harness a foundation model pretrained on over 10 million recordings, to evaluate ECG-derived age deviation as a cross-cohort biomarker of CVD burden. A predictive model, trained exclusively on healthy subjects, achieved accurate age prediction. Diseased subjects exhibited significant positive age acceleration across multiple categories, with structural and ischemic heart diseases showing the largest effects. External validation in a hospital-based cohort (n=160,493) confirmed that age acceleration independently predicts all-cause mortality, with the strongest prognostic value in patients under 65 years. Furthermore, we demonstrated that disease discrimination and mortality prediction are preserved across 6-lead and single-lead configurations, supporting potential deployment in wearable or mobile devices. Our analysis also revealed a striking morphological confound from the complete left bundle branch block, leading us to propose absolute age deviation as a more robust, universal risk marker. These findings establish ECG-derived biological age deviation as a highly generalizable and clinically actionable biomarker for assessing cardiovascular risk. We have also developed a web application at https://bioinformatics.mdc-berlin.de/ECGage that allows users to easily test our framework.
Ronan, R.; Tarabanis, C.; Chinitz, L.; Jankelosn, L.
Show abstract
Existing deep learning algorithms for electrocardiogram (ECG) classification rely on supervised training approaches requiring large volumes of reliably labeled data. This limits their applicability to rare cardiac diseases like Brugada syndrome (BrS), often lacking accurately labeled ECG examples. To address labeled data constraints and the resulting limitations of supervised training approaches, we developed a novel deep learning model for BrS ECG classification using the Variance-Invariance-Covariance Regularization (VICReg) architecture for self-supervised pre-training. The VICReg model outperformed a state-of-the-art neural network in all calculated metrics, achieving an area under the receiver operating and precision-recall curves of 0.88 and 0.82, respectively. We used the VICReg model to identify missed BrS cases and hence refine the previously underestimated institutional BrS prevalence and patient outcomes. Our results provide a novel approach to rare cardiac disease identification and challenge existing BrS prevalence estimates offering a framework for other rare cardiac conditions.
Nezamabadi, K.; Sivalokanathan, S.; Lee, J. W.; Tanriverdi, T.; Chen, M.; Lu, D.-Y.; Abraham, J.; Sardaripour, N.; Li, P.; Mousavi, P.; Abraham, M. R.
Show abstract
Left ventricular (LV) scar is a risk factor for sudden cardiac death and heart failure in hypertrophic cardiomyopathy (HCM). LV scar is frequent in HCM and evolves over time. Hence there is a need for LV scar detection and longitudinal monitoring. The current gold standard for LV scar detection is late gadolinium enhancement (LGE) on magnetic resonance imaging (MRI), which is limited by high cost and susceptibility to artifacts from implanted defibrillators. We introduce XplainScar, the first explainable machine learning method for LV scar detection and localization in HCM, using 12-lead electrocardiogram (ECG) data, which is not influenced by implanted devices. We use 500 patients from the JH-HCM Registry for model development, and 248 patients from the UCSF-HCM-Registry for validation. XplainScar combines unsupervised and self-supervised ECG representation learning, resulting in high precision (90%), sensitivity (95%), specificity (80%) and F1-score (90%) for scar detection in the basal, mid, and apical LV myocardium, with a processing time of <1 minute per 10 patients. Basal LV scar prediction by XplainScar is dominated by QRS features, and mid/apical LV scar by T wave features. XplainScar generalizes well to the held-out test UCSF data, with 88% precision, 90% sensitivity, 78% specificity, and F1-score of 89%. In summary, XplainScar demonstrates good performance for LV scar detection, and provides ECG signatures of basal, mid, and apical LV scar in HCM. XplainScar is publicly available at https://github.com/KasraNezamabadi/XplainScar
Crystal, O.; Farina, J. M. M.; Scalia, I. G.; Ayoub, C.; Park, H.-B.; Kim, K. A.; Arsanjani, R.; Lester, S. J.; Banerjee, I.
Show abstract
BackgroundAccurate assessment of left ventricular outflow tract (LVOT) gradients is critical for hypertrophic cardiomyopathy (HCM) management, yet Doppler-based measurements are technically demanding and require expertise. ObjectiveTo develop a multi-view deep learning model capable of classifying LVOT obstruction (> 20mmHg) using routine 2D echocardiographic windows without reliance on Doppler imaging. MethodsWe trained and externally validated a cross-attention-based video-to-video fusion framework that integrated EchoPrime-derived video representations from three standard transthoracic echocardiographic views to classify LVOT gradients. ResultsTraining was performed on a derivation cohort (N = 1833) from a tertiary care system in the United States, with model performance evaluated on an internal held-out test set (N = 275) and a Korean external validation cohort (N = 46). Single-view baselines showed limited discrimination (external AUROCs 0.47-0.70). Conversely, domain-specific foundational model (EchoPrime) achieved superior single-view performance (AUROCs 0.75-0.80 internal; 0.79-0.83 external), highlighting the importance of echo-specific pretraining and temporal modeling. The proposed multi-view fusion further enhanced predictive performance, with the late fusion model reaching an AUROC of 0.84 on the external cohort with significant population-shift. ConclusionsThese results suggest LVOT physiology is encoded in routine 2D imaging and can be leveraged for clinically relevant gradient classification without Doppler input- proposed AI-guided strategy demonstrates substantial cost savings compared with the screen-all approach. By integrating complementary spatial-temporal information across multiple views, our approach generalizes robustly across populations and may enable real-time decision support, extend LVOT assessment to portable or resource-limited settings, and complement Doppler-based evaluation for longitudinal HCM management.
Oikonomou, E. K.; Holste, G.; Coppi, A.; McNamara, R. L.; Nadkarni, G.; Baloescu, C.; Krumholz, H.; Wang, Z.; Khera, R.
Show abstract
BackgroundPoint-of-care ultrasonography (POCUS) enables cardiac imaging at the bedside and in communities but is limited by abbreviated protocols and variation in quality. We developed and tested artificial intelligence (AI) models to automate the detection of underdiagnosed cardiomyopathies from cardiac POCUS. MethodsIn a development set of 290,245 transthoracic echocardiographic videos across the Yale-New Haven Health System (YNHHS), we used augmentation approaches and a customized loss function weighted for view quality to derive a POCUS-adapted, multi-label, video-based convolutional neural network (CNN) that discriminates HCM (hypertrophic cardiomyopathy) and ATTR-CM (transthyretin amyloid cardiomyopathy) from controls without known disease. We evaluated the final model across independent, internal and external, retrospective cohorts of individuals who underwent cardiac POCUS across YNHHS and Mount Sinai Health System (MSHS) emergency departments (EDs) (2011-2024) to prioritize key views and validate the diagnostic and prognostic performance of single-view screening protocols. FindingsWe identified 33,127 patients (median age 61 [IQR: 45-75] years, n=17,276 [52{middle dot}2%] female) at YNHHS and 5,624 (57 [IQR: 39-71] years, n=1,953 [34{middle dot}7%] female) at MSHS with 78,054 and 13,796 eligible cardiac POCUS videos, respectively. An AI-enabled single-view screening approach successfully discriminated HCM (AUROC of 0{middle dot}90 [YNHHS] & 0{middle dot}89 [MSHS]) and ATTR-CM (YNHHS: AUROC of 0{middle dot}92 [YNHHS] & 0{middle dot}99 [MSHS]). In YNHHS, 40 (58{middle dot}0%) HCM and 23 (47{middle dot}9%) ATTR-CM cases had a positive screen at median of 2{middle dot}1 [IQR: 0{middle dot}9-4{middle dot}5] and 1{middle dot}9 [IQR: 1{middle dot}0-3{middle dot}4] years before clinical diagnosis. Moreover, among 24,448 participants without known cardiomyopathy followed over 2{middle dot}2 [IQR: 1{middle dot}1-5{middle dot}8] years, AI-POCUS probabilities in the highest (vs lowest) quintile for HCM and ATTR-CM conferred a 15% (adj.HR 1{middle dot}15 [95%CI: 1{middle dot}02-1{middle dot}29]) and 39% (adj.HR 1{middle dot}39 [95%CI: 1{middle dot}22-1{middle dot}59]) higher age- and sex-adjusted mortality risk, respectively. InterpretationWe developed and validated an AI framework that enables scalable, opportunistic screening of treatable cardiomyopathies wherever POCUS is used. FundingNational Heart, Lung and Blood Institute, Doris Duke Charitable Foundation, BridgeBio Research in Context Evidence before this studyPoint-of-care ultrasonography (POCUS) can support clinical decision-making at the point-of-care as a direct extension of the physical exam. POCUS has benefited from the increasing availability of portable and smartphone-adapted probes and even artificial intelligence (AI) solutions that can assist novices in acquiring basic views. However, the diagnostic and prognostic inference from POCUS acquisitions is often limited by the short acquisition duration, suboptimal scanning conditions, and limited experience in identifying subtle pathology that goes beyond the acute indication for the study. Recent solutions have shown the potential of AI-augmented phenotyping in identifying traditionally under-diagnosed cardiomyopathies on standard transthoracic echocardiograms performed by expert operators with strict protocols. However, these are not optimized for opportunistic screening using videos derived from typically lower-quality POCUS studies. Given the widespread use of POCUS across communities, ambulatory clinics, emergency departments (ED), and inpatient settings, there is an opportunity to leverage this technology for diagnostic and prognostic inference, especially for traditionally under-recognized cardiomyopathies, such as hypertrophic cardiomyopathy (HCM) or transthyretin amyloid cardiomyopathy (ATTR-CM) which may benefit from timely referral for specialized care. Added value of this studyWe present a multi-label, view-agnostic, video-based convolutional neural network adapted for POCUS use, which can reliably discriminate cases of ATTR-CM and HCM versus controls across more than 90,000 unique POCUS videos acquired over a decade across EDs affiliated with two large and diverse health systems. The model benefits from customized training that emphasizes low-quality acquisitions as well as off-axis, non-traditional views, outperforming view-specific algorithms and approaching the performance of standard TTE algorithms using single POCUS videos as the sole input. We further provide evidence that among reported controls, higher probabilities for HCM or ATTR-CM-like phenotypes are associated with worse long-term survival, suggesting possible under-diagnosis with prognostic implications. Finally, among confirmed cases with previously available POCUS imaging, positive AI-POCUS screens were seen at median of 2 years before eventual confirmatory testing, highlighting an untapped potential for timely diagnosis through opportunistic screening. Implications of all available evidenceWe define an AI framework with excellent performance in the automated detection of underdiagnosed yet treatable cardiomyopathies. This framework may enable scalable screening, detecting these disorders years before their clinical recognition, thus improving the diagnostic and prognostic inference of POCUS imaging in clinical practice.
Jeong, S.; Moon, I.; Jeon, J.; Jeong, D.; Lee, J.; kim, J.; Lee, S.-A.; Jang, Y.; Yoon, Y. E.; Chang, H.-J.
Show abstract
BackgroundPericardial disease spans a wide spectrum from small effusions to life-threatening tamponade or constriction. Transthoracic echocardiography (TTE) is the main diagnostic tool, but its interpretation is limited by operator dependence and incomplete functional assessment. Existing deep learning (DL) models focus mainly on effusion detection, lacking broader evaluation. MethodsWe developed a DL-based framework that performs sequential assessment of pericardial disease: (1) morphological features, including effusion amount (normal/small/moderate/large) and pericardial thickening/adhesion (yes/no), from five B-mode views, and (2) hemodynamic significance (yes/no), incorporating Doppler and inferior vena cava measurements. The developmental dataset comprises 2,253 TTEs from multiple Korean institutions (225 for internal testing), and the independent external test set consists of 274 TTEs. ResultsIn the internal test set, diagnostic accuracy was 81.8-97.3% for effusion, 91.6% for thickening/adhesion, and 86.2% for hemodynamic significance. External test set accuracy was 80.3-94.2%, 94.5%, and 85.5%, respectively. Area under the receiver operating curves (AUROCs) for the three tasks was 0.92-0.99, 0.90, and 0.79 internally, and 0.95-0.98, 0.85, and 0.76 externally. Sensitivity for thickening/adhesion and hemodynamic significance improved from 66.7% to 77.3%, and 68.8% to 80.8%, respectively, when poor image quality were excluded. Similar performance gains were observed in subgroups with complete target views and a higher number of available video clips. ConclusionsThis study presents the first DL-based TTE model for broader pericardial disease evaluation, integrating morphological with supportive functional assessments. The proposed framework demonstrated strong generalizability and aligned with the real-world diagnostic workflow. However, caution is warranted when interpreting results under suboptimal imaging conditions.
Holste, G.; Oikonomou, E. K.; Wang, Z.; Khera, R.
Show abstract
ImportanceEchocardiography is a cornerstone of cardiovascular care but relies on expert interpretation and manual reporting from a series of videos. We propose an artificial intelligence (AI) system, PanEcho, to automate echocardiogram interpretation with multi-task deep learning. ObjectiveTo develop and evaluate the accuracy of PanEcho on a comprehensive set of 39 echocardiographic labels and measurements on transthoracic echocardiography (TTE). Design, Setting, and ParticipantsThis study represents the development and retrospective, multi-site validation of an AI system. PanEcho was developed using a sample of TTE studies conducted at Yale-New Haven Health System (YNHHS) hospitals and clinics from January 2016-June 2022 during routine care. The trained model was internally validated in a temporally distinct YNHHS cohort from July-December 2022, externally validated across four diverse external cohorts, and made publicly available. Main Outcomes and MeasuresThe primary outcome was the area under the receiver operating characteristic curve (AUC) for diagnostic classification tasks and mean absolute error (MAE) for parameter estimation tasks, comparing AI predictions with the assessment of the interpreting cardiologist. ResultsThis study included 1.2 million echocardiographic videos from 32,265 TTE studies of 24,405 patients across YNHHS hospitals and clinics. PanEcho performed 18 diagnostic classification tasks with a median AUC of 0.91 (IQR: 0.88-0.93) and estimated 21 echocardiographic parameters with a median normalized MAE of 0.13 (0.10-0.18) in internal validation. For instance, the model accurately estimated left ventricular (LV) ejection fraction (MAE: 4.2% internal; 4.5% external) and detected moderate or higher LV systolic dysfunction (AUC: 0.98 internal; 0.99 external), RV systolic dysfunction (0.93 internal; 0.94 external), and severe aortic stenosis (0.98 internal; 1.00 external). PanEcho maintained excellent performance in limited imaging protocols, performing 15 diagnosis tasks with 0.91 median AUC (IQR: 0.87-0.94) in an abbreviated TTE cohort and 14 tasks with 0.85 median AUC (0.77-0.87) on real-world point-of-care ultrasound acquisitions by non-experts from YNHHS emergency departments. Conclusions and RelevanceWe report an AI system that automatically interprets echocardiograms, maintaining high accuracy across geography and time from complete and limited studies. PanEcho may be used as an adjunct reader in echocardiography labs or rapid AI-enabled screening tool in point-of-care settings. KEY POINTSO_ST_ABSQuestionC_ST_ABSCan artificial intelligence (AI) fully automate echocardiogram interpretation? FindingsWe report the development and validation of an automated AI system for echocardiogram analysis, called PanEcho, that performed 18 diagnostic classification tasks with a median area under the receiver operating characteristic curve (AUC) of 0.91 and 21 echocardiographic parameter estimation tasks with a median normalized mean absolute error (MAE) of 0.14. MeaningAn AI system can automate complete echocardiogram interpretation with high accuracy, potentially accelerating workflows and enabling rapid cardiovascular health screening in point-of-care settings with limited access to trained experts.
Lopez-Gutierrez, P.; Morales-Galan, A.; Galian-Gay, L.; Garrido-Oliver, J.; Dux-Santoy, L.; Craig, N.; Ye, Z.; Alegret, J. M.; Bermejo, J.; Calvo-Iglesias, F.; Ferrer, E.; Mendez, I.; Robledo Carmona, J. M.; Sanchez-Sanchez, V.; Saura, D.; Sevilla, T.; Foley, T.; Cuellar-Calabria, H.; Ferreira-Gonzalez, I.; Michelena, H. I.; Dweck, M. R.; Evangelista, A.; Teixido-Tura, G.; Rodriguez-Palomares, J. F.; Guala, A.
Show abstract
BackgroundAortic valve calcification (AVC), as measured by gold-standard computed tomography (CT) Agatston score, provides an anatomic assessment of aortic stenosis (AS) severity and is a key predictor of AS progression and need for valve replacement. AVC detection and quantification from transthoracic echocardiography (TTE) could expand AS early diagnosis and risk stratification, currently limited by CT availability and radiation exposure. MethodsA multi-view video-based deep learning framework was developed using 1166 TTE aortic valve videos from 187 TTE studies acquired in 110 patients with available AVC score by CT. EchoAVC architecture includes a feature extraction model followed by a quality-aware model that aggregates video-level information to obtain patient-level predictions for AVC detection and quantification (score). The framework was validated internally with 173 TTE studies from 86 patients across seven centres, and externally using 430 TTE studies from 280 patients across four different centres. The associations between EchoAVC estimations and AS severity, progression, and need for aortic valve replacement were examined. ResultsEchoAVC demonstrated excellent performance in AVC detection (AUROC 0.98, accuracy 94.4%), and quantification (R = 0.64) in the external multi-centre testing set. EchoAVC score was correlated with echocardiographic AS severity descriptors, including aortic valve mean pressure gradient ({rho} = 0.749) and peak velocity ({rho} = 0.757), and predicted future increase in mean pressure gradient ({rho} = 0.382), peak velocity ({rho} = 0.433) and calcium score by CT ({rho} = 0.650). In 361 patients followed for a median of 3.8 years, 139 underwent aortic valve replacement. Baseline presence and extent of AVC as predicted by EchoAVC showed strong risk-stratification power for aortic valve replacement, remarkably in line with those obtained by CT, and incremental over TTE AS descriptors. EchoAVC was further tested in routine clinical practice images, confirming strong associations with AS severity and progression, including stratification for incident AS in previously unaffected individuals (p<0.001). ConclusionsEchoAVC enables accurate and non-invasive detection and quantification of AVC, offering substantial diagnostic and prognostic value for aortic stenosis progression and need for valve replacement. This technique holds promise as a scalable tool for early detection and clinical management of aortic valve stenosis. Clinical PerspectiveO_ST_ABSWhat Is New?C_ST_ABS* Aortic valve calcification can be detected and quantified on standard, two-dimensional transthoracic echocardiography * EchoAVC score is associated with echocardiographic metrics of aortic stenosis severity * EchoAVC predicts progressive increase in aortic stenosis severity and need for aortic valve replacement What Are the Clinical Implications?* This openly available deep learning framework can be used to estimate aortic valve calcium in patients at risk of or presenting aortic stenosis to predict aortic stenosis progression * EchoAVC may help identify patients likely to have high aortic valve calcium and prioritize them for confirmatory CT assessment, thereby supporting earlier identification of calcific aortic valve disease in clinical pathways.