EchoFM: A Pre-training and Fine-tuning Framework for Echocardiogram Videos Vision Foundation Model
Zhang, Z.; Wu, Q.; Ding, S.; Wang, X.; Ye, J.
Show abstract
BackgroundEchocardiograms provide essential insights into cardiac health, yet their complex, multidimensional data poses significant challenges for analysis and interpretation. Existing deep learning models for echocardiogram analysis often rely heavily on supervised training, which limits their generalizability and robustness across different datasets and clinical environments. ObjectiveTo develop and evaluate Echo-Vision-FM (Echocardiogram video Vision Foundation Model), a self-supervised video learning framework designed to pre-train a video encoder on large-scale, unlabeled echocardiogram data. Echo-Vision-FM aims to produce robust and transferable video representations, improving downstream performance across diverse echocardiogram datasets and clinical conditions. MethodsThe proposed framework employs advanced self-supervised video learning through a masked auto-encoding technique, which compresses segments of video data and reconstructs the full video by masking non-overlapping video patches. An asymmetric encoder-decoder architecture underpins this approach. To further enhance the learned representations, we introduce STF-Net, a Spatial-Temporal Fusion Net, which integrates spatial and temporal correlations from the video representations. We pre-trained Echo-Vision-FM using the MIMIC-IV-ECHO dataset and fine-tuned it across multiple downstream datasets for specific clinical tasks, including morphological value estimation and the diagnosis of heart function and diseases. ResultsEcho-Vision-FM achieved superior performance in classifying left ventricular ejection fraction (LVEF), with an accuracy of 0.905, an F1 score of 0.941, and an AUC of 0.931. In regression tasks, Echo-Vision-FM outperformed state-of-the-art models, achieving a mean absolute error (MAE) of 3.87% and an r2 of 0.825 for LVEF prediction. The model also demonstrated significant improvements in estimating end-systolic and end-diastolic volumes, with r2 values of 0.782 and 0.742, respectively. Incorporating STF-Net further enhanced performance across all tasks. ConclusionOur results demonstrate that large-scale self-supervised video learning on echocardiogram data enables the extraction of transferable and clinically relevant features, surpassing existing methods. The Echo-Vision-FM framework, particularly with the inclusion of STF-Net, significantly improves the extraction of spatiotemporal features, resulting in enhanced predictive accuracy for a range of cardiac parameters. Echo-Vision-FM offers a scalable and effective solution for echocardiogram analysis, with promising applications in clinical diagnostics and research.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 94%
- Multiple Instance Learning Framework can Facilitate Explainability in Murmur Detection 94%
- A recurrent neural network and parallel hidden Markov model algorithm to segment and detect heart murmurs in phonocardiograms 93%
Similar papers in this journal
- Dual-Field Microvascular Segmentation: Hemodynamically-Consistent Attention Learning for Retinal Vasculature Mapping 94%
- pathCLIP: Detection of Genes and Gene Relations from Biological Pathway Figures through Image-Text Contrastive Learning 92%
- A Transformer-Based Model Trained on Large Scale Claims Data for Prediction of Severe COVID-19 Disease Progression 91%
Similar papers in this journal
- Creating a computer assisted ICD coding system: performance metric choice and use of the ICD hierarchy 92%
- A methodology of phenotyping ICU patients from EHR data: high-fidelity, personalized, and interpretable phenotypes estimation 91%
- ARCH: Large-scale Knowledge Graph via Aggregated Narrative Codified Health Records Analysis 91%
Similar papers in this journal
- GLAPAL-H: Global, Local, And Parts Aware Learner for Hydrocephalus Infection Diagnosis in Low-Field MRI 94%
- Saak Transform-Based Machine Learning for Light-Sheet Imaging of Cardiac Trabeculation 94%
- Maximum Classifier Discrepancy Generative Adversarial Network for Jointly Harmonizing Scanner Effects and Improving Reproducibility of Downstream Tasks 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.