Back

Self-supervised Learning for Chest CT - Training Strategies and Effect on Downstream Applications

Tariq, A.; Patel, B.; Banerjee, I.

2024-02-05 radiology and imaging
10.1101/2024.02.01.24302144 medRxiv
Show abstract

Self-supervised pretraining can reduce the amount of labeled training data needed by pre-learning fundamental visual characteristics of the medical imaging data. In this study, we investigate several self-supervised training strategies for chest computed tomography exams and their effects of downstream applications. we bench-mark five well-known self-supervision strategies (masked image region prediction, next slice prediction, rotation prediction, flip prediction and denoising) on 15M chest CT slices collected from four sites of Mayo Clinic enterprise. These models were evaluated for two downstream tasks on public datasets; pulmonary embolism (PE) detection (classification) and lung nodule segmentation. Image embeddings generated by these models were also evaluated for prediction of patient age, race, and gender to study inherent biases in models understanding of chest CT exams. Use of pretraining weights, especially masked regions prediction based weights, improved performance and reduced computational effort needed for downstream tasks compared to task-specific state-of-the-art (SOTA) models. Performance improvement for PE detection was observed for training dataset sizes as large as [Formula] with maximum gain of 5% over SOTA. Segmentation model initialized with pretraining weights learned twice as fast as randomly initialized model. While gender and age predictors built using self-supervised training weights showed no performance improvement over randomly initialized predictors, the race predictor experienced a 10% performance boost when using self-supervised training weights. We released models and weights under open-source academic license. These models can then be finetuned with limited task-specific annotated data for a variety of downstream imaging tasks thus accelerating research in biomedical imaging informatics.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Journal of Medical Imaging
11 papers in training set
Top 0.1%
13.3%
2
Medical Physics
14 papers in training set
Top 0.1%
13.1%
3
Scientific Reports
3612 papers in training set
Top 4%
9.9%
4
PLOS ONE
5266 papers in training set
Top 24%
6.8%
5
PLOS Digital Health
106 papers in training set
Top 1%
4.9%
6
European Radiology
15 papers in training set
Top 0.2%
4.5%
50% of probability mass above
7
Medical Image Analysis
35 papers in training set
Top 0.2%
4.4%
8
IEEE Transactions on Medical Imaging
21 papers in training set
Top 0.1%
4.1%
9
IEEE Access
35 papers in training set
Top 0.3%
4.1%
10
Expert Systems with Applications
11 papers in training set
Top 0.1%
1.9%
11
Nature Communications
5641 papers in training set
Top 44%
1.8%
12
Diagnostics
50 papers in training set
Top 1%
1.8%
13
Physics in Medicine & Biology
18 papers in training set
Top 0.2%
1.7%
14
npj Digital Medicine
118 papers in training set
Top 2%
1.5%
15
Frontiers in Computational Neuroscience
60 papers in training set
Top 0.8%
1.5%
16
PLOS Computational Biology
1863 papers in training set
Top 15%
1.5%
17
European Heart Journal - Digital Health
18 papers in training set
Top 0.7%
1.4%
18
Nature Machine Intelligence
70 papers in training set
Top 2%
1.4%
19
Computers in Biology and Medicine
128 papers in training set
Top 3%
1.4%
20
GigaScience
212 papers in training set
Top 3%
1.1%
21
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.4%
1.1%
22
Journal of the American Medical Informatics Association
71 papers in training set
Top 2%
1.0%
23
BMC Medical Informatics and Decision Making
43 papers in training set
Top 2%
1.0%
24
Biomedical Optics Express
95 papers in training set
Top 1.0%
0.9%
25
Photoacoustics
12 papers in training set
Top 0.3%
0.6%
26
NeuroImage
903 papers in training set
Top 6%
0.6%
27
Frontiers in Neuroinformatics
41 papers in training set
Top 0.7%
0.6%