Back

Pan-cancer subclonal mutation analysis of 7,827 tumors predicts clinical outcome

Jiang, Y.; Montierth, M. D.; Yu, K.; Ji, S.; Guo, S.; Tran, Q.; Liu, X.; Shin, S. J.; Cao, S.; Li, R.; Tang, Y.; Lesluyes, T.; Kopetz, S.; Ajani, J.; Msaouel, P.; Subudhi, S. K.; Aparicio, A.; Sharma, P.; Shen, J. P.; Sood, A. K.; Tarabichi, M.; Wang, J. R.; Kimmel, M.; Loo, P. V.; Zhu, H.; Wang, W.

2024-07-06 genomics
10.1101/2024.07.03.601939 bioRxiv
Show abstract

Intra-tumor heterogeneity is characterized by a diverse population of tumor clones and subclones which are important drivers of tumor evolution and therapeutic response. However, accurate subclonal reconstruction at scale remains challenging. We developed a machine learning tool, CliPP, and surveyed 9,972 tumors from 32 cancer types. We found that high subclonal mutation load (sML), the fraction of subclonal single nucleotide variants (SNVs) to all SNVs in the coding region, was prognostic of survival (progression free survival or overall survival) in 18 cancer types. In 14 cancers with low to moderate tumor mutation burden (TMB), high sML was associated with better prognosis. In immunotherapy trials for 42 metastatic prostate cancer (mCRPC), high sML was predictive of favorable response to ipilimumab and associated with increased CD8+ T-cell infiltration and decreased macrophage population. A validation using 613 whole-genomes of esophageal adenocarcinoma confirms the favorable effect of high sML and the observed tumor-associated macrophage. Our study identifies sML as a key feature of cancer, suggesting a biphasic relationship between evolutionary dynamics and differential immune environments. Finally, sML may serve as an orthogonal approach to identify likely responders of immune checkpoint blockade in low to moderate TMB tumors.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.