Back

Screening of Cellular Senescence Related Genes as Biomarkers and Therapeutic Targets for Glioblastoma by Integrated Machine Learning

Xue, T.; Hu, Y.; Lyu, H.; Xie, T.

2025-08-24 cancer biology
10.1101/2025.08.19.670890 bioRxiv
Show abstract

Glioblastoma (GBM) is an aggressive brain tumor with limited prognostic biomarkers and therapeutic targets. This study applied an integrated machine learning (IML) framework to discover cellular senescence (CS)-related gene biomarkers and candidate therapeutic targets in GBM. Using Gene Expression Omnibus (GEO) training and validation cohorts, 113 machine learning models across 11 algorithms were integrated to pinpoint CS-associated gene signatures. Differential expression analysis combined with overlap of a CellAge senescence gene set yielded 129 CS-related differentially expressed genes (DEGs). Functional enrichment of these DEGs highlighted pathways related to cellular senescence and cell cycle regulation (Gene Ontology and KEGG). A multivariate classifier constructed via stepwise generalized linear modeling (GLM) and LASSO achieved high diagnostic performance (area under the ROC curve 0.92 in training, each independent validation set is over 0.85). Seven top-ranked genes from this model were validated, with TGF{beta}I emerging as the most robust biomarker (AUC > 0.85) and its elevated expression was associated with significantly shorter overall survival. Spatial transcriptomics and single-cell RNA sequencing localized TGF{beta}I expression to tumor cell clusters harboring high copy number variation (CNV) burdens. Immune microenvironment profiling linked TGF{beta}I expression with increased macrophage infiltration. Finally, single-cell gene set enrichment (scGSEA) and AUCell analyses indicated enrichment of ECM-receptor interaction signaling in TGF{beta}I-expressing cells. In summary, IML combined with spatial and single-cell transcriptomics identified TGF{beta}I as a potent CS-related biomarker and a promising therapeutic target in GBM.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.