Back

Vision-Language Foundation Models Do Not Transfer to Medical Imaging Classification: A Negative Result on Chest X-ray Diagnosis

Fisher, G. R.

2025-12-08 radiology and imaging
10.64898/2025.12.06.25341759 medRxiv
Show abstract

Vision-language models (VLMs) pretrained on web-scale data have achieved remarkable performance across diverse tasks, leading to widespread adoption in industry. A natural question is whether these powerful representations transfer to specialized medical imaging domains, and whether domain-specific medical pretraining improves transfer. We tested these hypotheses using two VLMs on the NIH ChestX-ray14 benchmark: Qwen2.5-VL (pretrained on web data) and BiomedCLIP (pretrained on 15 million PubMed biomedical image-text pairs). Both models dramatically underperformed compared to convolutional neural networks (CNNs) with ImageNet pretraining. Across 5 random seeds, the best VLM achieved F1=0.196 {+/-} 0.004 versus a CNN baseline of F1=0.811. Domain-specific pretraining provided marginal improvement: BiomedCLIPs frozen encoder achieved F1=0.161 {+/-} 0.001 versus Qwens F1=0.124 (+30%), but this remains clinically inadequate. Fine-tuning both models led to catastrophic overfitting, with sensitivity collapsing from >65% to <36% as the models learned to predict "no disease" for all inputs. These results demonstrate that neither general-purpose nor medical-specific vision-language pretraining produces features suitable for dense multi-label medical imaging classification. For chest X-ray diagnosis, traditional CNNs with ImageNet pretraining remain substantially more effective than VLM-based approaches.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
38.3%
2
npj Digital Medicine
118 papers in training set
Top 0.7%
9.4%
3
Scientific Reports
3612 papers in training set
Top 14%
6.1%
50% of probability mass above
4
Nature Communications
5641 papers in training set
Top 28%
5.4%
5
Patterns
78 papers in training set
Top 0.4%
3.9%
6
Medical Image Analysis
35 papers in training set
Top 0.2%
3.9%
7
PLOS ONE
5266 papers in training set
Top 36%
3.4%
8
European Heart Journal - Digital Health
18 papers in training set
Top 0.4%
3.1%
9
Nature Medicine
125 papers in training set
Top 0.9%
2.7%
10
Journal of the American Medical Informatics Association
71 papers in training set
Top 1%
2.7%
11
The Lancet Digital Health
25 papers in training set
Top 0.3%
1.7%
12
npj Precision Oncology
53 papers in training set
Top 1%
1.4%
13
Communications Medicine
113 papers in training set
Top 3%
1.3%
14
eBioMedicine
183 papers in training set
Top 4%
1.1%
15
Advanced Intelligent Systems
11 papers in training set
Top 0.3%
0.8%
16
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.8%
0.8%
17
Journal of Medical Imaging
11 papers in training set
Top 0.4%
0.8%
18
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 46%
0.6%
19
GigaScience
212 papers in training set
Top 6%
0.6%
20
JAMIA Open
42 papers in training set
Top 2%
0.6%
21
Frontiers in Medicine
120 papers in training set
Top 5%
0.6%
22
European Radiology
15 papers in training set
Top 0.7%
0.6%
23
Journal of Medical Internet Research
87 papers in training set
Top 3%
0.6%
24
Modern Pathology
22 papers in training set
Top 0.5%
0.6%
25
PeerJ
308 papers in training set
Top 14%
0.6%