Back

A unified 12-lead ECG-language model for interpretation and clinical-endpoint prediction

Banerjee, R.; Beaulieu, J.; Dostie, N.; Mondesert, B.; Ahuja, R.; Nolin-Lapalme, A.; Jabbour, G.; Srikanth, S. S.; Marquis-Gravel, G.; Tastet, O.; Sowa, A.; Cadrin-Tourigny, J.; Delfrate, J.; Avram, R.

2026-07-24 cardiovascular medicine
10.64898/2026.07.22.26358591 medRxiv
Show abstract

Automated electrocardiogram (ECG) interpretation has advanced, yet most systems remain narrow classifiers that emit fixed labels rather than the narratives or endpoint-specific answers clinicians need. Generative approaches could instead produce rich narratives, but are constrained by the gap between continuous biosignals and discrete language tokens. Here we present DeepECG-Tok, which reframes ECG interpretation as a unified instruction-following problem. A residual vector-quantization tokenizer (QINCo) maps 12-lead waveforms to language-model-compatible tokens. Its frozen embeddings achieved a macro-averaged area under the receiver operating characteristic curve (AUROC) of 0.96 for 77-condition classification, outperforming supervised and self-supervised baselines. Aligned with a large language model, a single instruction-tuned model performed ECG interpretation, structured reporting and clinical endpoint prediction, including left ventricular ejection fraction, structural heart disease and atrial fibrillation risk, using 7.27 million question answer pairs. Endpoint classifiers using the frozen tokenizer transferred without retraining to four external cohorts, retaining AUROCs of 0.88--0.90. Evaluated end to end using an ontology-grounded large language model as a judge, which achieved a mean agreement (Cohen's K) of 0.82 against two cardiologists, the unified instruction-tuned model scored 0.71 internally and 0.50--0.53 in the same external cohorts. In a blinded reader study, board-certified cardiologists and residents rated its free-text reports comparably to reference clinician reports, with a paired win-tie-loss distribution of 32:35:33 and a forced-choice preference of 0.53 among decided cases, with no significant difference between the model and reference reports in either comparison. These findings establish discrete ECG tokenization as a foundation for general-purpose models that generate clinically useful interpretations and answer diverse questions directly from cardiac waveforms.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.