Deep Learning Features Encode Interpretable Morphologies within Histological Images
Foroughi Pour, A.; White, B. S.; Park, J.; Sheridan, T. B.; Chuang, J. H.
Show abstract
Convolutional neural networks (CNNs) are revolutionizing digital pathology by enabling machine learning-based classification of a variety of phenotypes from hematoxylin and eosin (H&E) whole slide images (WSIs), but the interpretation of CNNs remains difficult. Most studies have considered interpretability in a post hoc fashion, e.g. by presenting example regions with strongly predicted class labels. However, such an approach does not explain the biological features that contribute to correct predictions. To address this problem, here we investigate the interpretability of H&E-derived CNN features (the feature weights in the final layer of a transfer-learning-based architecture), which we show can be construed as abstract morphological genes ("mones") with strong independent associations to biological phenotypes. We observe that many mones are specific to individual cancer types, while others are found in multiple cancers especially from related tissue types. We also observe that mone-mone correlations are strong and robustly preserved across related cancers. Importantly, linear mone-based classifiers can very accurately separate 38 distinct classes (19 tumor types and their adjacent normals, AUC=97.1% {+/-} 2.8% for each class prediction), and linear classifiers are also highly effective for universal tumor detection (AUC=99.2% {+/-} 0.12%). This linearity provides evidence that individual mones or correlated mone clusters may be associated with interpretable histopathological features or other patient characteristics. In particular, the statistical similarity of mones to gene expression values allows integrative mone analysis via expression-based bioinformatics approaches. We observe strong correlations between individual mones and individual gene expression values, notably mones associated with collagen gene expression in ovarian cancer. Mone-expression comparisons also indicate that immunoglobulin expression can be identified using mones in colon adenocarcinoma and that immune activity can be identified across multiple cancer types, and we verify these findings by expert histopathological review. Our work demonstrates that mones provide a morphological H&E decomposition that can be effectively associated with diverse phenotypes, analogous to the interpretability of transcription via gene expression values.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MorphLink: Bridging Cell Morphological Behaviors and Molecular Dynamics in Multi-modal Spatial Omics 96%
- Machine learning-based tissue of origin classification for cancer of unknown primary diagnostics using genome-wide mutation features 96%
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 96%
Similar papers in this journal
Similar papers in this journal
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 98%
- Deciphering tumor ecosystems at super-resolution from spatial transcriptomics with TESLA 96%
- Uncovering the spatial landscape of molecular interactions within the tumor microenvironment through latent spaces 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.