Benchmarking pathology foundation models for non-neoplastic pathology in the placenta
Peng, Z.; Ayad, M. A.; Jing, Y.; Chou, T.; Cooper, L. A. D.; Goldstein, J. A.
Show abstract
Machine learning (ML) applications within diagnostic histopathology have been extremely successful. While many successful models have been built using general-purpose models trained largely on everyday objects, there is a recent trend toward pathology-specific foundation models, trained using histopathology images. Pathology foundation models show strong performance on cancer detection and subtyping, grading, and predicting molecular diagnoses. However, we have noticed lacunae in the testing of foundation models. Nearly all the benchmarks used to test them are focused on cancer. Neoplasia is an important pathologic mechanism and key concern in much of clinical pathology, but it represents one of many pathologic bases of disease. Non-neoplastic pathology dominates findings in the placenta, a critical organ in human development, as well as a specimen commonly encountered in clinical practice. Very little to none of the data used in training pathology foundation models is placenta. Thus, placental pathology is doubly out of distribution, representing a useful challenge for foundation models. We developed benchmarks for estimation of gestational age, classifying normal tissue, identifying inflammation in the umbilical cord and membranes, and in classification of macroscopic lesions including villous infarction, intervillous thrombus, and perivillous fibrin deposition. We tested 5 pathology foundation models and 4 non-pathology models for each benchmark in tasks including zero-shot K-nearest neighbor classification and regression, content-based image retrieval, supervised regression, and whole-slide attention-based multiple instance learning. In each task, the best performing model was a pathology foundation model. However, the gap between pathology and non-pathology models was diminished in tasks related to inflammation or those in which a supervised task was performed using model embeddings. Performance was comparable among pathology foundation models. Among non-pathology models, ResNet consistently performed worse, while models from the present decade showed better performance. Future work could examine the impact of incorporating placental data into foundation model training.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving Pre-eclampsia Risk Prediction by Modeling Individualized Pregnancy Trajectories Derived from Routinely Collected Electronic Medical Record Data 94%
- Clinical Knowledge Extraction via Sparse Embedding Regression (KESER) with Multi-Center Large Scale Electronic Health Record Data 93%
- Zero-shot Interpretable Phenotyping of Postpartum Hemorrhage Using Large Language Models 92%
Similar papers in this journal
- Interpretable multimodal deep learning for real-time pan-tissue pan-disease pathology search on social media 94%
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 93%
- Artificial intelligence-driven morphology-based enrichment of malignant cells from body fluid 93%
Similar papers in this journal
- A user-friendly tool for cloud-based whole slide image segmentation, with examples from renal histopathology 94%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 92%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 92%
Similar papers in this journal
- Weakly-Supervised Tumor Purity Prediction FromFrozen H&E Stained Slides 94%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 93%
- Integrative deep learning analysis improves colon adenocarcinoma patient stratification at risk for mortality 93%
Similar papers in this journal
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 93%
- Development and validation of AI-based pre-screening of large bowel biopsies 93%
- Multicenter Validation of a Machine Learning Algorithm for Diagnosing Pediatric Patients with Multisystem Inflammatory Syndrome and Kawasaki Disease 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.