A Multi-Task Deep Learning Model for Pediatric Echocardiography Analysis
Cho, J.; Mathur, M.; Kaur, D.; Duda, M.; Dahlan, A.; Krishnan, A.; Leipzig, M.; Shad, R.; Gonzalez, A. K.; Logan, J.; Seidman, C.; Fong, R.; Kumar, A.; Zakka, C.; Langlotz, C. P.; Jolley, M. A.; Hiesinger, W.
Show abstract
BackgroundCongenital heart defects afflict nearly 1% of all births worldwide. While deep learning algorithms have shown significant promise in automating and improving adult echocardiography analysis, similar progress has not been observed in pediatric echocardiography. Specifically, existing pediatric-based models are limited to single tasks and specific echocardiographic views. To address this, we introduce EchoAI-Peds, the first multi-task deep learning model for pediatric echocardiography. Our model was developed using the most comprehensive set of pediatric labels to date and is designed to integrate information from multiple echocardiographic views simultaneously. MethodsA video-based vision transformer was trained to simultaneously detect 28 congenital heart defects, structural and functional abnormalities, repairs, and interventions directly from complete pediatric echocardiography studies with multiple videos. During inference, our model integrates information from all available views to produce unified study-level predictions. Our model was developed using over 700,000 videos derived from more than 11,000 studies at Stanford Medicine. Model efficacy was tested on an internal held-out dataset. In addition, model generalizability was tested on a spatially and temporally distinct patient cohort at the Childrens Hospital of Philadelphia. ResultsOur model achieved macro-averaged AUROC values of 0.91 (95% CI: 0.90-0.92) and 0.89 (95% CI: 0.88-0.90) on the internal and external test sets, respectively. Moreover, our model significantly outperformed adult-based echocardiography foundation models trained on substantially larger datasets (p < 0.001). Finally, our model demonstrated robust performance across patient age, patient sex, and studies with varying number of videos. ConclusionsOur findings demonstrate the remarkable potential for multi-task deep learning models to aid the interpretation of pediatric echocardiograms. In addition, our results underscore the need for models that are specifically tailored to pediatric populations.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multimodal deep learning enhances diagnostic precision in left ventricular hypertrophy 96%
- Simple Models Versus Deep Learning in Detecting Low Ejection Fraction From The Electrocardiogram 94%
- Development and Multinational Validation of an Ensemble Deep Learning Algorithm for Detecting and Predicting Structural Heart Disease Using Noisy Single-lead Electrocardiograms 94%
Similar papers in this journal
- Designing a computer-assisted diagnosis system for cardiomegaly detection and radiology report generation 92%
- Detecting QT prolongation From a Single-lead ECG With Deep Learning 91%
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 90%
Similar papers in this journal
- Biometric Contrastive Learning for Data-Efficient Deep Learning from Electrocardiographic Images 96%
- A Comparative Analysis of Privacy-Preserving Large Language Models For Automated Echocardiography Report Analysis 95%
- ENRICHing Medical Imaging Training Sets Enables More Efficient Machine Learning 94%
Similar papers in this journal
- Opportunistic Assessment of Ischemic Heart Disease Risk Using Abdominopelvic Computed Tomography and Medical Record Data: a Multimodal Explainable Artificial Intelligence Approach 93%
- Explaining Deep Neural Networks for Knowledge Discovery in Electrocardiogram Analysis 92%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.