Automated Diagnostic Reports from Images of Electrocardiograms at the Point-of-Care
Khunte, A.; Sangha, V.; Oikonomou, E. K.; Dhingra, L. S.; Aminorroaya, A.; Coppi, A.; Vasisht Shankar, S.; Mortazavi, B. J.; Bhatt, D. L.; Krumholz, H. M.; Nadkarni, G.; Vaid, A.; Khera, R.
Show abstract
BackgroundTimely and accurate assessment of electrocardiograms (ECGs) is crucial for diagnosing, triaging, and clinically managing patients. Current workflows rely on computerized ECG interpretation tools built into ECG signal acquisition systems, which use rule-based algorithms that are unreliable and frequently not available in low-resource settings. We developed and validated a format-independent vision encoder-decoder model - ECG-GPT - that can generate free-text, expert-level interpretations directly from 12-lead ECG images. MethodsUsing 12-lead ECGs and their corresponding diagnosis statements collected at the Yale-New Haven Health System (YNHHS) between 2000 and 2022, we developed a vision-text transformer model to generate interpretation statements from images of ECGs. Using structured clinical assessment, semantic similarity, and conventional natural language generation metrics, we validated ECG-GPT across 7 geographically distinct health settings. These include (1) 3 large and diverse US health systems, (2) consecutive ECGs from a central reading system in Minas Gerais, Brazil, (3) the prospective cohort study, UK Biobank, (4) a Germany-based, publicly available repository, PTB-XL, and (5) a community hospital in Missouri. ResultsOverall, 2.9 million ECGs were used for model development. The model performed well in clinical assessment across 26 extracted labels: for atrial fibrillation, sinus tachycardia, sinus bradycardia, premature atrial contractions, and premature ventricular contractions, AUROCs and AUPRCs ranged from 0.80-0.95 and 0.50-0.86, respectively. For left bundle branch block, right bundle branch block, first degree atrioventricular block, left anterior fascicular block, and left posterior fascicular block, AUROCs and AUPRCs ranged from 0.88-0.96 and 0.23-0.86, respectively. Across all 26 conditions, diagnostic accuracy ranged between 0.93-0.99. ECG-GPT identified the full context of the diagnosis statements with allied conditions. It had a median pairwise cosine similarity of 0.90 (IQR 0.83-0.97), significantly greater than the median baseline similarity of 0.73 (IQR 0.67-0.78, p<0.001). This separation between median pairwise and baseline similarity remained consistent across all 26 condition-specific subsets. The results were comparable across external validation sites. ConclusionsWe developed and extensively validated a vision encoder-decoder model that generates expert-level interpretations from ECG images. This represents a scalable and accessible strategy for automated ECG analysis, especially in low-resource settings. CLINICAL PERSPECTIVEO_ST_ABSWhat is New?C_ST_ABSO_LIECG-GPT is a vision encoder-decoder model capable of generating full-text ECG interpretations directly from ECG images, regardless of layout or format. C_LIO_LIThe model was trained on over 2.7 million ECGs and externally validated across 3.8 million additional ECGs from demographically and geographically diverse populations. C_LI What are the clinical implications?O_LIECG-GPT enables automated, expert-level ECG interpretation directly from images, eliminating the need for signal data or device integration. C_LIO_LIThe model demonstrates consistent performance across diverse patient populations, ECG formats, and care settings. C_LIO_LIThis scalable, image-based approach may expand access to accurate ECG interpretation in low-resource settings. C_LI
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Biometric Contrastive Learning for Data-Efficient Deep Learning from Electrocardiographic Images 98%
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 95%
- ENRICHing Medical Imaging Training Sets Enables More Efficient Machine Learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.