Back

Path2Space: An AI Approach for Cancer Biomarker Discovery Via Histopathology Inferred Spatial Transcriptomics

Shulman, E. D.; Campagnolo, E. M.; Lodha, R.; Cantore, T.; Hu, T.; Nasrallah, M.; Hoang, D.-T.; Aldape, K.; Ruppin, E.

2024-10-18 cancer biology
10.1101/2024.10.16.618609 bioRxiv
Show abstract

Spatial transcriptomics (ST) is transforming our understanding of tumor heterogeneity by enabling high-resolution, location-specific mapping of gene expression across tumors and their microenvironment. However, the associated high cost of the assay has limited cohort size and hence large-scale biomarker discovery. Here we present Path2Space, a deep learning approach that predicts spatial gene expression directly from histopathology slides. Trained on substantial breast cancer ST data, it robustly predicts the spatial expression of over 4,300 genes in independent validations, markedly outperforming existing ST predictors. Path2Space additionally accurately infers cell-type abundances in the tumor microenvironment (TME) based on the inferred ST data. Applied to more than a thousand breast tumor histopathology slides from the TCGA, Path2Space characterizes their TME on an unprecedented scale and identifies three new spatially-grounded breast cancer subgroups with distinct survival rates. Path2Space-inferred TME landscapes enable more accurate predictions of patients response to chemotherapy and trastuzumab directly from H&E slides than those obtained by existing established sequencing-based biomarkers. Path2Space thus offers a transformative, fast and cost-effective approach to robustly delineate the TME directly from their histopathology slides, facilitating the development of spatially-grounded biomarkers to advance precision oncology.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.