Back

Histology and spatial transcriptomic integration revealed infiltration zone with specific cell composition as a prognostic hotspot in glioblastoma

Fidon, L.; Lubrano, M.; Hoffmann, C.; Baena, E.; Thiriez, C.; Ayestas, L. C.; Cornish, A. J.; Madissoon, E.; Ducret, V.; Esposito, C.; Durand, E. Y.; Schmauch, B.; Pronier, E.; Romagnoni, A.; Espin-Perez, A.; MOSAIC consortium, ; Eckstein, M.; Maragkou, T.; Shelan, M.; Youssef, A.; Watson, S.; Sanson, M.; Bielle, F.

2025-10-08 cancer biology
10.1101/2025.10.08.681087 bioRxiv
Show abstract

BackgroundGlioblastoma (GBM), the most aggressive primary brain tumor, has a median survival of approximately 15 months. Twenty percent of patients survive beyond three years, but known clinical factors like age, performance status, resection extent, and MGMT promoter methylation status do not fully explain the observed outcomes. ObjectiveOur objective was to identify novel histology derived biomarkers associated with end-of-spectrum overall survival (OS) to provide novel biological insight with a translational potential. MethodsWe analyzed a total of 748 GBM patients from 3 different cohorts, uniquely enriched in long survivors (n=98 with overall survival (OS) > 5y including n=196 with OS[&ge;]3y), with clinical data and H&E slides obtained from the primary tumor at baseline. We propose an interpretable machine learning (ML) methodology for the discovery of histological biomarkers. Our method learned to segment each H&E slide into three distinct regions associated with long-term survival, short-term survival, and non-informative tissue. We characterized these regions by integrating unsupervised learning, nuclei segmentation, blood vessels detection, pathologist annotations, and multimodal data including spatial transcriptomics from n=31 patients of the GBM MOSAIC dataset to discover fully interpretable biomarkers. ResultsOur OS prediction model using histology and clinical data as input achieved an area under the curve (AUC) of 0.85 for the classification of patients between OS<2 and OS[&ge;]3y in external cohort validation, outperforming significantly models trained on clinical data or on histology alone (AUC of 0.76; 0.73, respectively). Two novel biomarkers were predicting poor survival: the presence of regions of lowly infiltrated white matter enriched in malignant cells with a mesenchymal-like phenotype, and lower levels of angiogenesis associated with higher hypoxia response in the main tumor regions. We also found that a subtype of immunosuppressive tumor macrophages - defined by high PLIN2 expression and lipid accumulation- is consistently enriched in histological areas predictive of poor prognosis. ConclusionOur interpretable ML methodology identified a novel prognostic impact of biological processes and cell types according to distinct tumor regions of GBM. These results pave the way for spatially-informed biomarkers to improve risk stratification and for personalized spatially-targeted therapeutic strategies. Key highlightsO_LIOur ML model identified histological biomarkers predicting prognosis independently from known clinical factors C_LIO_LIThe region of lowly infiltrated white matter enriched in malignant cells including a mesenchymal-like phenotype is predictive of poor prognosis C_LIO_LIAngiogenesis is increased in areas predictive of long survival in main non-necrotic tumor regions. C_LIO_LIThe subtype of macrophages expressing PLIN2 and associated with increased lipid metabolism was associated with poor prognosis in all GBM regions. C_LI Highlights O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=116 SRC="FIGDIR/small/681087v4_ufig1.gif" ALT="Figure 1"> View larger version (43K): org.highwire.dtl.DTLVardef@45e619org.highwire.dtl.DTLVardef@105a4b0org.highwire.dtl.DTLVardef@17f257dorg.highwire.dtl.DTLVardef@7669b4_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.