Self-Supervised Pretraining Enables High-Performance Chest X-Ray Interpretation Across Clinical Distributions
Iyer, N. S.; Gulati, A.; Banerjee, O.; Loge, C.; Farhat, M.; Saenz, A.; Rajpurkar, P.
Show abstract
Chest X-rays (CXRs) are a rich source of information for physicians - essential for disease diagnosis and treatment selection. Recent deep learning models aim to alleviate strain on medical resources and improve patient care by automating the detection of diseases from CXRs. However, shortages of labeled CXRs can pose a serious challenge when training models. Currently, models are generally pretrained on ImageNet, but they often need to then be finetuned on hundreds of thousands of labeled CXRs to achieve high performance. Therefore, the current approach to model development is not viable on tasks with only a small amount of labeled data. An emerging method for reducing reliance on large amounts of labeled data is self-supervised learning (SSL), which uses unlabeled CXR datasets to automatically learn features that can be leveraged for downstream interpretation tasks. In this work, we investigated whether self-supervised pretraining methods could outperform traditional ImageNet pretraining for chest X-ray interpretation. We found that SSL-pretrained models outperformed ImageNet-pretrained models on thirteen different datasets representing high diversity in geographies, clinical settings, and prediction tasks. We thus show that SSL on unlabeled CXR data is a promising pretraining approach for a wide variety of CXR interpretation tasks, enabling a shift away from costly labeled datasets.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 96%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 95%
- Enhancing Fairness in Disease Prediction by Optimizing Multiple Domain Adversarial Networks 94%
Similar papers in this journal
Similar papers in this journal
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 95%
- An ML prediction model based on clinical parameters and automated CT scan features for COVID-19 patients 95%
- ARA: accurate, reliable and active histopathological image classification framework with Bayesian deep learning 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.