Multi-modal foundation model with whole-slide attention enables transferrable digital pathology at single-cell resolution
Wu, Q.; Gong, Q.; Yuan, L.; Li, Z.; Ashenberg, O.; Chen, F.; Xavier, R.; Uhler, C.
Show abstract
Paired histopathology and spatial transcriptomics data are advancing our understanding of tissue biology and disease, but modeling both modalities at single-cell resolution while mapping local and distal cell-cell interdependencies remains computationally prohibitive. Here we introduce TissueFormer, a framework for pretraining foundation models with linear rather than quadratic computational complexity, overcoming a long-standing barrier to modeling long-range dependencies at scale. Trained on over 17 million image-expression pairs from 1.2K tissue slides, TissueFormer excels at predicting spatial gene expression from histology images at cellular resolution and scales to diagnostic tasks at the cell, region, and slide levels. Additionally, by identifying both long and short-range cell-cell interdependencies, our model enables the generation of testable hypotheses about disease mechanisms and staging, as demonstrated in lung fibrosis and breast cancer.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DGAT: A Dual-Graph Attention Network for Inferring Spatial Protein Landscapes from Transcriptomics 96%
- CAPTAIN: A multimodal foundation model pretrained on co-assayed single-cell RNA and protein 96%
- Multi-modal Diffusion Model with Dual-Cross-Attention for Multi-Omics Data Generation and Translation 95%
Similar papers in this journal
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 97%
- Identifying maximally informative signal-aware representations of single-cell data using the Information Bottleneck 94%
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 94%
Similar papers in this journal
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 94%
- Multi-V-Stain: Multiplexed Virtual Staining of Histopathology Whole-Slide Images 94%
- AI for radiographic COVID-19 detection selects shortcuts over signal 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.