Real-World Benchmarking and Validation of Foundation Model Transformers for Endometrial Cancer Subtyping from Histopathology
Wagner, V. M.; Cosgrove, C. M.; Chen, S. J.; Griffin, D. T.; Samuelson, M. I.; Goodheart, M. J.; Gonzalez Bosquet, J.
Show abstract
PurposeTo evaluate whether open-source histopathology foundation model pipelines, paired with attention-based multiple instance learning (MIL), can accurately classify molecular subtypes of endometrial cancer (EC) from whole-slide images (WSIs) and maintain performance in a real-world, independent cohort. MethodsWe assembled a public discovery cohort of 815 patients (1,195 WSIs) from The Cancer Genome Atlas and Clinical Proteomic Tumor Analysis Consortium, and an independent external cohort of 720 patients (1,357 WSIs) with molecular subtyping determined by mismatch repair immunohistochemistry plus TP53 and POLE sequencing. Four ImageNet-pretrained convolutional neural networks (CNNs) and six open-source foundation encoders using two MIL aggregation strategies (TransMIL and CLAM) were benchmarked within the STAMP pipeline. Models were trained with five-fold cross-validation and evaluated on an independent cohort. Macro-area under the receiver operating characteristic curve (AUC) was the primary outcome. ResultsIn cross-validation, foundation models outperformed CNNs (macro-AUC 0.799-0.860 vs 0.715-0.829). The best configuration (Virchow2 with CLAM) achieved macro-AUC 0.860 (95%CI, 0.839-0.880), macro-F1 score 0.607, and balanced accuracy 0.647. External validation showed substantial degradation for CNNs, while foundation models retained higher discrimination (macro-AUC 0.667-0.780). UNI2 with CLAM had the highest external macro-AUC (0.780), and Virchow2 with CLAM had the best balanced accuracy (0.525). Subtype-level AUCs for UNI2 with CLAM were highest for p53abn (0.851). ConclusionsOpen-source foundation model pipelines with attention-based MIL can deliver accurate and generalizable molecular subtyping of EC directly from WSIs. These models outperform CNNs in real-world validation, supporting their potential as scalable, cost-effective tools to guide precision oncology and triage confirmatory molecular testing.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial intelligence-based histopathology image analysis identifies a novel subset of endometrial cancers with distinct genomic features and unfavourable outcome 97%
- The Impact of Digital Histopathology Batch Effect on Deep Learning Model Accuracy and Bias 96%
- Integrative ensemble modelling of cetuximab sensitivity in colorectal cancer PDXs 95%
Similar papers in this journal
- Multi-resolution deep learning characterizestertiary lymphoid structures in solid tumors 95%
- LUNAR: A Deep Learning Model to Predict Glioma Recurrence Using Integrated Genomic and Clinical Data 95%
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 93%
Similar papers in this journal
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 96%
- Weakly-Supervised Tumor Purity Prediction FromFrozen H&E Stained Slides 96%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 93%
Similar papers in this journal
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 94%
- Spatial Transcriptomics Inferred from Pathology Whole-Slide Images Links Tumor Heterogeneity to Survival in Breast and Lung Cancer 94%
- Construction and optimization of multi-platform precision pathways for precision medicine 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.