Back

Structural semantic evolutionary distance (SSED) unifies the selection of cancer driver genes across macroevolution and tumorigenesis.

Ji, H.; Zhang, Q.; Wu, J.; Tian, F.; Wang, X.; Chai, X.; Chang, J.; Zheng, M.; Li, X.; Zhang, H.-M.

2025-12-19 genetics
10.64898/2025.12.17.694808 bioRxiv
Show abstract

The non-random, site-specific enrichment of somatic mutations in cancer driver genes (CDGs) suggests their emergence is governed by underlying evolutionary constraints. However, quantitative methods to define these constraints and their underlying principles remain underexplored. To address this, we introduce the Structural Semantic Evolutionary Distance (SSED), a metric leveraging the pretrained ESM-3 protein language model to quantify evolutionary divergence within a unified structural semantic space. Our analysis demonstrates that CDGs are subject to persistent structural semantic constraints across species, tolerating a significantly narrower range of structural semantic changes during evolution compared to non-CDGs. Crucially, clinically observed oncogenic mutations follow this same principle, favoring minimal structural perturbation as shaped by long-term gene evolution. Such mutations maintain core protein function while conferring a capacity for immune evasion, thereby driving clonal expansion. Guided by this "evolutionary constraint" framework, we successfully predicted and experimentally validated a previously uncharacterized oncogenic mutation, KRAS R135L, in bronchial epithelial cells. Furthermore, clinical cohort analysis demonstrated that SSED acts as an independent predictor of response to immune checkpoint blockade, offering information orthogonal to tumor mutational burden (TMB). This study unifies the evolutionary principles governing CDGs across macroevolutionary and microevolutionary (tumorigenesis) timescales, elucidates the balance between structural adaptability, functional conservation, and immune pressure, and identifies a novel predictive biomarker for cancer immunotherapy.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.