Machine learning enabled prediction of digital biomarkers from whole slide histopathology images
McCaw, Z. R.; Shcherbina, A.; Shah, Y.; Huang, D.; Elliott, S.; Szabo, P. M.; Dulken, B.; Holland, S.; Tagari, P.; Light, D.; Koller, D.; Probert, C.
Show abstract
Current predictive biomarkers generally leverage technologies such as immunohis-tochemistry or genetic analysis, which may require specialized equipment, be time-intensive to deploy, or incur human error. In this paper, we present an alternative approach for the development and deployment of a class of predictive biomarkers, leveraging deep learning on digital images of hematoxylin and eosin (H&E)-stained biopsy samples to simultaneously predict a range of molecular factors that are relevant to treatment selection and response. Our framework begins with the training of a pan-solid tumor H&E foundation model, which can generate a universal featurization of H&E-stained tissue images. This featurization becomes the input to machine learning models that perform multi-target, pan-cancer imputation. For a set of 352 drug targets, we show the ability to predict with high accuracy: copy number amplifications, target RNA expression, and an RNA-derived "amplification signature" that captures the transcriptional consequences of an amplification event. We facilitate exploratory analyses by making broad predictions initially. Having identified the subset of biomarkers relevant to a patient population of interest, we develop specialized machine learning models, built on the same foundational featurization, which achieve even higher performance for key biomarkers in tumor types of interest. Moreover, our models are robust, generalizing with minimal loss of performance across different patient populations. By generating imputations from tile-level featurizations, we enable spatial overlays of molecular annotations on top of whole-slide images. These annotation maps provide a clear means of interpreting the histological correlates of our models predictions, and align with features identified by expert pathologist review. Overall, our work demonstrates a flexible and scalable framework for imputing molecular measurements from H&E, providing a generalizable approach to the development and deployment of predictive biomarkers for targeted therapeutics in cancer.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Image-Based Consensus Molecular Subtyping in Rectal Cancer Biopsies and Response to Neoadjuvant Chemoradiotherapy 95%
- Generalizing AI-driven Assessment of Immunohistochemistry across Immunostains and Cancer Types: A Universal Immunohistochemistry Analyzer 95%
- Deep learning inference of cell type-specific gene expression from breast tumor histopathology 95%
Similar papers in this journal
- Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features 96%
- Teacher-student collaborated multiple instance learning for pan-cancer PDL1 expression prediction from histopathology slides 95%
- Evolutionary signatures of human cancers revealed via genomic analysis of over 35,000 patients 95%
Similar papers in this journal
- Multi-resolution deep learning characterizestertiary lymphoid structures in solid tumors 94%
- Subpopulation-specific Machine Learning Prognosis for Underrepresented Patients with Double Prioritized Bias Correction 94%
- LUNAR: A Deep Learning Model to Predict Glioma Recurrence Using Integrated Genomic and Clinical Data 94%
Similar papers in this journal
Similar papers in this journal
- STPath: A Generative Foundation Model for Integrating Spatial Transcriptomics and Whole Slide Images 95%
- Biologically-informed deep neural networks provide quantitative assessment of intratumoral heterogeneity in post-treatment glioblastoma 94%
- High-Sensitivity Pan-Cancer AI Assessment of Lymph Node Metastasis via Uncertainty Quantification 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.