Back

Spatial mapping of immunosuppressive cancer-associated fibroblast gene signatures in H&E-stained images using additive multiple instance learning

Markey, M.; Kim, J.; Goldstein, Z.; Gerardin, Y.; Brosnan-Cashman, J.; Javed, S. A.; Juyal, D.; Padigela, H.; Yu, L.; Rahsepar, B.; Abel, J.; Hennek, S.; Khosla, A.; Taylor-Weiner, A.; Parmar, C.

2024-08-15 cancer biology
10.1101/2024.08.12.607604 bioRxiv
Show abstract

The relative abundance of cancer-associated fibroblast (CAF) subtypes influences a tumors response to treatment, especially immunotherapy. However, the extent to which the underlying tumor composition associates with CAF subtype-specific gene expression is unclear. Here, we describe an interpretable machine learning (ML) approach, additive multiple instance learning (aMIL), to predict bulk gene expression signatures from H&E-stained whole slide images (WSI), focusing on an immunosuppressive LRRC15+ CAF-enriched TGF{beta}-CAF signature. aMIL models accurately predicted TGF{beta}-CAF across various cancer types. Tissue regions contributing most highly to slide-level predictions of TGF{beta}-CAF were evaluated by ML models characterizing spatial distributions of diverse cell and tissue types, stromal subtypes, and nuclear morphology. In breast cancer, regions contributing most to TGF{beta}-CAF-high predictions ("excitatory") were localized to cancer stroma with high fibroblast density and mature collagen fibers. Regions contributing most to TGF{beta}-CAF-low predictions ("inhibitory") were localized to cancer epithelium and densely inflamed stroma. Fibroblast and lymphocyte nuclear morphology also differed between excitatory and inhibitory regions. Thus, aMIL enables a data-driven link between tissue phenotype and transcription, offering biological interpretability beyond typical black-box models.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.