A unified 12-lead ECG-language model for interpretation and clinical-endpoint prediction
Banerjee, R.; Beaulieu, J.; Dostie, N.; Mondesert, B.; Ahuja, R.; Nolin-Lapalme, A.; Jabbour, G.; Srikanth, S. S.; Marquis-Gravel, G.; Tastet, O.; Sowa, A.; Cadrin-Tourigny, J.; Delfrate, J.; Avram, R.
Show abstract
Automated electrocardiogram (ECG) interpretation has advanced, yet most systems remain narrow classifiers that emit fixed labels rather than the narratives or endpoint-specific answers clinicians need. Generative approaches could instead produce rich narratives, but are constrained by the gap between continuous biosignals and discrete language tokens. Here we present DeepECG-Tok, which reframes ECG interpretation as a unified instruction-following problem. A residual vector-quantization tokenizer (QINCo) maps 12-lead waveforms to language-model-compatible tokens. Its frozen embeddings achieved a macro-averaged area under the receiver operating characteristic curve (AUROC) of 0.96 for 77-condition classification, outperforming supervised and self-supervised baselines. Aligned with a large language model, a single instruction-tuned model performed ECG interpretation, structured reporting and clinical endpoint prediction, including left ventricular ejection fraction, structural heart disease and atrial fibrillation risk, using 7.27 million question answer pairs. Endpoint classifiers using the frozen tokenizer transferred without retraining to four external cohorts, retaining AUROCs of 0.88--0.90. Evaluated end to end using an ontology-grounded large language model as a judge, which achieved a mean agreement (Cohen's K) of 0.82 against two cardiologists, the unified instruction-tuned model scored 0.71 internally and 0.50--0.53 in the same external cohorts. In a blinded reader study, board-certified cardiologists and residents rated its free-text reports comparably to reference clinician reports, with a paired win-tie-loss distribution of 32:35:33 and a forced-choice preference of 0.53 among decided cases, with no significant difference between the model and reference reports in either comparison. These findings establish discrete ECG tokenization as a foundation for general-purpose models that generate clinically useful interpretations and answer diverse questions directly from cardiac waveforms.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Clinical and genetic associations of asymmetric apical and septal left ventricular hypertrophy 93%
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 93%
- Multimodal deep learning enhances diagnostic precision in left ventricular hypertrophy 91%
Similar papers in this journal
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 91%
- Polygenic score informed by genome-wide association studies of multiple ancestries and related traits improves risk prediction for coronary artery disease 91%
- Exome-by-phenome-wide rare variant gene burden association with electronic health record phenotypes 90%
Similar papers in this journal
Similar papers in this journal
- Cohort Design and Natural Language Processing to Reduce Bias in Electronic Health Records Research: The Community Care Cohort Project 94%
- Identification of Digital Twins to Guide Interpretable AI for Diagnosis and Prognosis in Heart Failure 93%
- Understanding the robustness of vision-language models to medical image artefacts 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.