Extended pre-training of histopathology foundation models uncovers co-existing breast cancer archetypes characterized by RNA splicing or TGF-β dysregulation
Fournier, L. M.; Haefliger, G.; Vernhes, A. E.; Jung, V.; Letovanec, I.; Frossard, P.; Vincent-Cuaz, C.; Luisier, R.
Show abstract
In recent years, histopathology foundation models (hFM) have rapidly advanced in size and complexity, achieving excellent performance in tasks such as cancer diagnosis and biomarker discovery. Here, we reveal novel capabilities of these models by specializing hFMs, originally trained on diverse tissue types, specifically to invasive tumor tissue. It enables unprecedented discrimination of visually similar yet molecularly distinct tumor regions, that were previously indistinguishable by baseline models, which eventually leads to uncovering new biological insights into breast cancer. Our contributions are threefold. First, to the best of our knowledge, this is the first study to systematically evaluate the biological concepts encoded within hFM representations across multiple scales. Second, we explore extended pre-training to identify optimal conditions that enhance the models ability to encode richer, tumor tissue-specific biological concepts. We show that this refinement strategy transforms generalist models into a specialist one capable of resolving subtle, recurrent tumor regions with distinct morphological and molecular identities, called tumor archetypes. Finally, leveraging this specialized model, we uncover two dominant tumor archetypes in invasive breast cancer characterized by distinct aberrant gene expression signatures, notably RNA metabolism dysregulation and TGF-{beta} signaling. Strikingly, these archetypes coexist within the same tumors as spatially distinct regions with varying densities and patterns, and are recurrent across patients, highlighting their universality and potential clinical relevance for patient stratification. Altogether our study demonstrates how extended pre-training of state-of-the-art hFM with specific tumor tissues can unlock rich molecular and morphological information encoded in H&E images. By providing a more accessible approach to investigating tumor heterogeneity, this work opens new avenues for precision oncology, using routine histopathology slides and low computational resources. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=128 SRC="FIGDIR/small/645192v1_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@38e660org.highwire.dtl.DTLVardef@19ce335org.highwire.dtl.DTLVardef@108e8e6org.highwire.dtl.DTLVardef@1f25bc0_HPS_FORMAT_FIGEXP M_FIG Graphical abstract C_FIG
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- STAIG: Spatial Transcriptomics Analysis via Image-Aided Graph Contrastive Learning for Domain Exploration and Alignment-Free Integration 97%
- Learning tissue representation by identification of persistent local patterns in spatial omics data 97%
- stLearn: integrating spatial location, tissue morphology and gene expression to find cell types, cell-cell interactions and spatial trajectories within undissociated tissues 96%
Similar papers in this journal
Similar papers in this journal
- Predicting MammaPrint Recurrence Risk from Breast Cancer Pathological Images Using a Weakly Supervised Transformer 97%
- Hierarchical prediction and perturbation of chromatin organization reveal how loop domains mediate higher-order architectures 94%
- Cancer-like fragmentomic characteristics of somatic variants in cell-free DNA 94%
Similar papers in this journal
- Predicting Gene Spatial Expression and Cancer Prognosis: An Integrated Graph and Image Deep Learning Approach Based on HE Slides 95%
- Cell states and neighborhoods in distinct clinical stages of primary and metastatic esophageal adenocarcinoma 95%
- TimiGP: inferring inter-cell functional interactions and clinical values in the tumor immune microenvironment through gene pairs 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.