Explainable Prototype Booster: Enhancing Latent Representations of Foundation Models for Gene Expression Prediction
Li, C.; Nguyen, Q.
Show abstract
Spatial transcriptomics (ST) is a cutting-edge technology that measures gene expression while preserving spatial context and generating pathology-grade tissue images. Although ST has enabled numerous discoveries and demonstrated a huge application potential in pathological diagnosis and prognosis, the technology remains time-consuming and costly. The ability to predict gene markers of cancer from histological H&E-stained tissue images can overcome these technological barriers to open new horizons for precision and personalised pathology. Recently, foundation models have demonstrated improvements in generating general-purpose embeddings of H&E-images. However, these improved representations are not optimized for gene expression prediction and lack task-specific adaptability. To address this limitation, we propose the Explainable Prototype Booster (EP-Booster), which incorporates biological prior knowledge to guide the construction and training of learnable prototypes for embedding refinement, thereby improving gene expression prediction. Importantly, model predictions are inherently interpretable through pathway-level attributions associated with the prototypes. Extensive experiments across multiple datasets, cancer types, and spatial transcriptomics platforms demonstrate that EP-Booster consistently outperforms existing methods. Moreover, EP-Booster can be integrated with diverse foundation models to enhance task-specific representations, thereby improving predictive performance and biological interpretability in clinically relevant applications, including cancer biomarker prediction, survival analysis, and drug response prediction.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting Spatially Resolved Gene Expression via Tissue Morphology using Adaptive Spatial GNNs 97%
- TUGDA: Task uncertainty guided domain adaptation for robust generalization of cancer drug response prediction from in vitro to in vivo settings 96%
- Adversarial Deconfounding Autoencoder for Learning Robust Gene Expression Embeddings 96%
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 97%
- Features fusion or not: harnessing multiple pathological foundation models using Meta-Encoder for downstream tasks fine-tuning 96%
- SpaIM: Single-cell Spatial Transcriptomics Imputation via Style Transfer 96%
Similar papers in this journal
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 97%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 96%
- Novel multi-omics deconfounding variational autoencoders can obtain meaningful disease subtyping 96%
Similar papers in this journal
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 96%
- COSIME: Cooperative multi-view integration with Scalable and Interpretable Model Explainer 96%
- An interpretable deep learning framework for genome-informed precision oncology 95%
Similar papers in this journal
- Highly Accurate Cancer Phenotype Prediction with AKLIMATE, a Stacked Kernel Learner Integrating Multimodal Genomic Data and Pathway Knowledge 96%
- A Deep Learning approach for time-consistent cell cycle phase prediction from microscopy data 95%
- A variational autoencoder trained with priors from canonical pathways increases the interpretability of transcriptome data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.