Cracks in the Foundation: How Data-Hungry and Sensitive to Domain Shift are Vision Foundation Models for Computational Pathology?
Bonn, S.; Zimmermann, M.; Sauter, G.; Bengtsson, E.; Huber, T. B.; Baumbach, J.; Lennartz, M.; Fuhlert, P.; Witte, A.
Show abstract
BackgroundVision Foundation Models (VFM) have emerged as a promising approach for computational pathology, offering scalable feature representations that may reduce labelled-data requirements and improve robustness to variation in tissue preparation and digitisation. However, VFM decoder and dataset size requirements as well as the performance under real-world domain shifts remain unclear. MethodsWe evaluated six contemporary VFMs on a protocol-variant Prostate Cancer (PCa) dataset comprising 37 683 tissue microarray spot images from 10 412 patients. The dataset includes six controlled domain shifts arising from differences in staining duration, section thickness, scanner type, and sampling location. Two clinically relevant downstream tasks were examined: ISUP grading and 5-year relapse prediction. We compared two decoder architectures, quantified dataset-size requirements using a saturation analysis (45-5727 samples), and assessed cross-domain robustness using out-of-domain test sets. FindingsLarger VFMs consistently outperformed smaller models in peak accuracy and robustness metrics. Contrary to expectations of data efficiency, all models showed strong dependence on training-set size, requiring at least 1000 samples to approach stable results. All VFMs showed notable degradation under protocol-level domain shifts, with performance reductions of 4 to 13 percentage points in both cancer grading and relapse prediction, although larger models exhibited somewhat greater robustness. Furthermore, KNN-based probing performed substantially worse than a decoder-based approach across all architectures. InterpretationsDespite their strong representational capacity, current VFMs do not yet provide reliable domain generalisation or data-efficient performance in computational pathology. Decoder design remains essential, and substantial amounts of labelled data are still required to achieve clinically meaningful accuracy. Further advances in pre-training strategies, decoder architectures, and domain adaptation methods will be crucial for translating VFMs into robust clinical tools. Research in contextO_ST_ABSEvidence before this studyC_ST_ABSThis study focuses on pathology foundation models, which offer promising improvements in performance, data requirements, and robustness to domain shifts for computational pathology. To identify relevant studies, we searched in Google Scholar for research published before 1 April 2025. We searched for studies introducing novel foundation models trained on pathology images, or reviews comparing those models in terms of performance or robustness. The search terms used were computational pathology, benchmarking, review and foundation model, as well combinations of these terms. We found that many studies focus on increasing the complexity of pathology foundation models while using increasingly extensive and heterogeneous pre-training datasets. Various benchmarking studies demonstrate the superior performance and robustness of more recent and larger foundation models. However, these studies have limitations in their evaluation datasets. Either they cover only a domain shift due to a different scanner device, or they have small sample sizes. We also identified a research gap regarding the requirement for large datasets to train a decoder based on a pathology foundation model for a specific downstream task. Added value of this studyThe goal of this study was to evaluate the necessity of large downstream task datasets and the domain shift robustness of multiple pathology foundation models. For this purpose, we used our internal protocol-variant prostate cancer dataset, which provides a controlled evaluation setup as multiple domain shift types have been intentionally and separately introduced for different sub-datasets. Our saturation analysis revealed that at least 1000 samples were necessary to achieve good performance. Furthermore, our findings show that none of the evaluated foundation models are robust against all of our domain shifts, though larger models generally perform better. Implications of all the available evidenceThis study reveals that increasing the capacity of pathology foundation models improves performance and robustness. However, we demonstrated that all models exhibit some degree of performance degradation for certain domain shifts and require substantial datasets for training on downstream tasks. These limitations demonstrate that pathology foundation models do not fully address the issues of robustness and data requirements.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.