Cross-country generalizability of foundation models for cervical cancer screenings on H&E whole slide images
Paulikat, M.; Bosch, C.; Aswolinskiy, W.; Caixeta Borges, I.; Nauschuette, L.; Aichmueller, C.; Schmidt, D.; Bussmann, H.; Kalteis, S.; Zapukhlyak, M.; von Knebel Doeberitz, M.; Kloor, M.
Show abstract
Accurate grading of cervical biopsies on Hematoxylin and Eosin (H&E) stained whole slide images (WSIs) is essential for distinguishing high grade lesions from low grade changes, yet this process is subject to considerable inter-observer variability. In this study, we evaluate a foundation model-based multiple instance learning (MIL) pipeline for binary high-grade squamous intraepithelial lesion (HSIL) detection on H&E stained WSIs. We benchmark our Athena foundation model against four state-of-the-art pathology foundation models: H-optimus-0, Hibou-L, Midnight-12k and Virchow, across datasets from five different countries: Portugal, Cambodia, Germany, Poland and Scotland. Athena achieved the highest mean area under the curve (AUC) (0.931) with the lowest cross-country variability (STD = 0.022). Furthermore, we compared the model's diagnostic performance to that of trained pathologists on a dataset with p16-confirmed ground truth. Our model improved sensitivity from 84% to 95% while maintaining comparable specificity (85% vs. 84%). Failure analysis revealed that the model's errors were concentrated at the diagnostic boundary between low-grade and high-grade lesions, whereas pathologists' errors spanned a broader range of misclassifications. These findings show the potential of foundation models for cervical cancer screenings worldwide.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 97%
- Built to last? Reproducibility and Reusability of Deep Learning Algorithms in Computational Pathology 96%
- Clinical-Grade Validation of an Autofluorescence Virtual Staining System with Human Experts and a Deep Learning System for Prostate Cancer 95%
Similar papers in this journal
- Independent assessment of a deep learning system for lymph node metastasis detection on the Augmented Reality Microscope 95%
- Using an Anomaly Detection Approach for the Segmentation of Colorectal Cancer Tumors in Whole Slide Images 94%
- Bladder Cancer Prognosis Using Deep Neural Networks and Histopathology Images 94%
Similar papers in this journal
- A Framework for Falsifiable Explanations of Machine Learning Models with an Application in Computational Pathology 94%
- ViFIT-assisted Histopathology: From H&E Style Standardization to Virtual Fiber Image Transformation 93%
- Public Covid-19 X-ray datasets and their impact on model bias - a systematic review of a significant problem 93%
Similar papers in this journal
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 95%
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 94%
- PathProfiler: Automated Quality Assessment of Retrospective Histopathology Whole-Slide Image Cohorts by Artificial Intelligence, A Case Study for Prostate Cancer Research 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.