Back

DeepGraphMut: A graph-based deep learning methodfor cancer prognosis using somatic mutation profile

Jose, A.; Srivastava, A.; Vinod, P. K.

2024-12-06 systems biology
10.1101/2024.12.03.626568 bioRxiv
Show abstract

Cancer remains a leading cause of morbidity and mortality worldwide. Despite advances in genomics, identifying clinically relevant subtypes of cancer remains challenging due to its complex and heterogeneous nature. In this work, we propose DeepGraphMut (DGM), a novel graph-based deep-learning pipeline that integrates somatic mutation data with protein-protein interaction (PPI) networks. By employing a graph autoencoder with a graph attention layer and a node-level attention decoder, DGM generates patient-specific clinically relevant encodings for unsupervised and supervised tasks. We demonstrate the effectiveness of DGM across 16 cancer types comprising of 7352 samples from The Cancer Genome Atlas (TCGA). Unsupervised clustering reveals distinct subtypes with significant survival differences in 11 cancer types. In supervised analysis using a Cox regression model, DGM demonstrates excellent performance in predicting survival outcomes, achieving a high concordance index (c-index) value in the range of 0.7 across most cancers, underscoring its robust predictive performance using only somatic mutation data. Furthermore, DGM outperforms its lightweight variant and the network-based stratification method in both unsupervised and supervised analyses. In summary, this study presents a promising approach for cancer subtype identification and prognosis, especially in resource-limited settings where multi-omics data may not be readily available. By leveraging the strengths of graph learning and network biology, DGM offers a valuable tool for advancing personalized medicine.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.