Back

ECG classification with convolutional neural networks demonstrates resilience to sex-imbalances in data

Schipaanboord, D. J. M.; van der Zalm, F. B. H.; van Es, R.; Vessies, M.; van de Leur, R. R.; Siegersma, K. R.; van der Harst, P.; den Ruijter, H. M.; Onland-Moret, N. C.; van Amsterdam, W. A. C.; on behalf of the IMPRESS consortium,

2025-09-02 cardiovascular medicine
10.1101/2025.08.31.25334560 medRxiv
Show abstract

BackgroundMany ECG-AI models have been developed to predict a wide range of cardiovascular outcomes. The underrepresentation of women in cardiovascular disease studies has raised concerns if these models are equally predictive in women as compared to men. We tested the effect of sex-imbalance in training datasets on predictive performance of ECG-AI models, investigating imbalance in representation (ratio women-to-men), as well as in outcome prevalence, and percentage of misclassification. MethodsWe used a dataset containing raw 12-lead ECGs (n = 474,006) of 181,755 individuals who visited the University Medical Center Utrecht at any of the non-cardiology departments between July 1997 and August 2023 and sampled a sex-balanced dataset (n = 165,156) including only one ECG per individual. Multiple deep convolutional neural networks were trained to predict four outcomes; left bundle branch block, Long QT Syndrome, left ventricular hypertrophy or ECGs classified as abnormal by a physician. Using subsampling, we simulated scenarios of sex-imbalance in representation (nscenario=5) for all outcomes and disease prevalence (nscenario=5), both representation and disease prevalence (nscenario=20) and disease misclassification (nscenario=7) for abnormal. Model performance was evaluated per scenario using area under the receiver operating characteristic curve (AUC) and smooth expected calibration error (smECE) for women and men separately. ResultsAcross all scenarios, the AUC remained stable, with small absolute differences between women and men for sex-imbalance in representation ({Delta}AUC: [0.002-0.025]), in disease prevalence ({Delta}AUC: [0.01-0.02]), in scenarios of both representation and disease prevalence ({Delta}AUC: [0.003-0.039]), and in outcome misclassification ({Delta}AUC: [0.007-0.077]). Only when disease prevalence in train and test data was sex-imbalanced, we observed differences in calibration error between sexes (max {Delta}smECE: 0.26), with similar patterns for women and men. ConclusionThe neural networks in this study demonstrated resilience to sex-imbalance in training ECG data. Graphical abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=139 SRC="FIGDIR/small/25334560v1_ufig1.gif" ALT="Figure 1"> View larger version (39K): org.highwire.dtl.DTLVardef@1731b2forg.highwire.dtl.DTLVardef@1fdc405org.highwire.dtl.DTLVardef@15063dborg.highwire.dtl.DTLVardef@cbe891_HPS_FORMAT_FIGEXP M_FIG C_FIG Graphical summary of the study methodology and results showing that ECG classification with convolutional neural networks is not sensitive to sex-imbalances in datasets. AUC = Area under the receiver operating curve; smECE = smooth expected calibration error. Created in BioRender. Meijer, I. (2025) https://BioRender.com/nxkwvoi.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.