Back

Cracks in the Foundation: How Data-Hungry and Sensitive to Domain Shift are Vision Foundation Models for Computational Pathology?

Bonn, S.; Zimmermann, M.; Sauter, G.; Bengtsson, E.; Huber, T. B.; Baumbach, J.; Lennartz, M.; Fuhlert, P.; Witte, A.

2026-01-06 urology
10.64898/2026.01.06.25342815 medRxiv
Show abstract

BackgroundVision Foundation Models (VFM) have emerged as a promising approach for computational pathology, offering scalable feature representations that may reduce labelled-data requirements and improve robustness to variation in tissue preparation and digitisation. However, VFM decoder and dataset size requirements as well as the performance under real-world domain shifts remain unclear. MethodsWe evaluated six contemporary VFMs on a protocol-variant Prostate Cancer (PCa) dataset comprising 37 683 tissue microarray spot images from 10 412 patients. The dataset includes six controlled domain shifts arising from differences in staining duration, section thickness, scanner type, and sampling location. Two clinically relevant downstream tasks were examined: ISUP grading and 5-year relapse prediction. We compared two decoder architectures, quantified dataset-size requirements using a saturation analysis (45-5727 samples), and assessed cross-domain robustness using out-of-domain test sets. FindingsLarger VFMs consistently outperformed smaller models in peak accuracy and robustness metrics. Contrary to expectations of data efficiency, all models showed strong dependence on training-set size, requiring at least 1000 samples to approach stable results. All VFMs showed notable degradation under protocol-level domain shifts, with performance reductions of 4 to 13 percentage points in both cancer grading and relapse prediction, although larger models exhibited somewhat greater robustness. Furthermore, KNN-based probing performed substantially worse than a decoder-based approach across all architectures. InterpretationsDespite their strong representational capacity, current VFMs do not yet provide reliable domain generalisation or data-efficient performance in computational pathology. Decoder design remains essential, and substantial amounts of labelled data are still required to achieve clinically meaningful accuracy. Further advances in pre-training strategies, decoder architectures, and domain adaptation methods will be crucial for translating VFMs into robust clinical tools. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSThis study focuses on pathology foundation models, which offer promising improvements in performance, data requirements, and robustness to domain shifts for computational pathology. To identify relevant studies, we searched in Google Scholar for research published before 1 April 2025. We searched for studies introducing novel foundation models trained on pathology images, or reviews comparing those models in terms of performance or robustness. The search terms used were computational pathology, benchmarking, review and foundation model, as well combinations of these terms. We found that many studies focus on increasing the complexity of pathology foundation models while using increasingly extensive and heterogeneous pre-training datasets. Various benchmarking studies demonstrate the superior performance and robustness of more recent and larger foundation models. However, these studies have limitations in their evaluation datasets. Either they cover only a domain shift due to a different scanner device, or they have small sample sizes. We also identified a research gap regarding the requirement for large datasets to train a decoder based on a pathology foundation model for a specific downstream task. Added value of this studyThe goal of this study was to evaluate the necessity of large downstream task datasets and the domain shift robustness of multiple pathology foundation models. For this purpose, we used our internal protocol-variant prostate cancer dataset, which provides a controlled evaluation setup as multiple domain shift types have been intentionally and separately introduced for different sub-datasets. Our saturation analysis revealed that at least 1000 samples were necessary to achieve good performance. Furthermore, our findings show that none of the evaluated foundation models are robust against all of our domain shifts, though larger models generally perform better. Implications of all the available evidenceThis study reveals that increasing the capacity of pathology foundation models improves performance and robustness. However, we demonstrated that all models exhibit some degree of performance degradation for certain domain shifts and require substantial datasets for training on downstream tasks. These limitations demonstrate that pathology foundation models do not fully address the issues of robustness and data requirements.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Journal of Medical Imaging
11 papers in training set
Top 0.1%
29.1%
2
Diagnostics
50 papers in training set
Top 0.2%
6.4%
3
npj Digital Medicine
118 papers in training set
Top 1.0%
5.7%
4
Scientific Reports
3612 papers in training set
Top 15%
5.6%
5
PLOS ONE
5266 papers in training set
Top 27%
5.6%
50% of probability mass above
6
Modern Pathology
22 papers in training set
Top 0.1%
4.9%
7
BMJ Open
601 papers in training set
Top 6%
3.3%
8
Cancers
213 papers in training set
Top 2%
2.5%
9
PLOS Computational Biology
1863 papers in training set
Top 12%
2.4%
10
Bioinformatics
1204 papers in training set
Top 6%
2.4%
11
Journal of Pathology Informatics
15 papers in training set
Top 0.1%
2.2%
12
PLOS Digital Health
106 papers in training set
Top 3%
1.7%
13
Biology Methods and Protocols
61 papers in training set
Top 0.8%
1.7%
14
Frontiers in Medicine
120 papers in training set
Top 2%
1.7%
15
Medical Physics
14 papers in training set
Top 0.4%
1.4%
16
Human Brain Mapping
329 papers in training set
Top 3%
1.2%
17
Sensors
43 papers in training set
Top 0.9%
1.2%
18
Radiotherapy and Oncology
19 papers in training set
Top 0.3%
1.2%
19
Clinical and Translational Radiation Oncology
10 papers in training set
Top 0.2%
1.2%
20
GigaScience
212 papers in training set
Top 3%
1.2%
21
BMC Medical Informatics and Decision Making
43 papers in training set
Top 1%
1.1%
22
Frontiers in Artificial Intelligence
20 papers in training set
Top 0.6%
1.1%
23
Biomedical Physics & Engineering Express
11 papers in training set
Top 0.3%
0.9%
24
The Prostate
11 papers in training set
Top 0.1%
0.9%