Essentiality, Protein-Protein Interactions and Evolutionary Properties are Key Predictors for Identifying Cancer Genes Using Machine Learning
Safadi, A.; Lovell, S. C.; Doig, A. J.
Show abstract
The identification of genes that may be linked to cancer is of great importance for the discovery of new drug targets. The rate at which cancer genes are being found experimentally is slow, however, due to the complexity of the identification and confirmation process, giving a narrow range of therapeutic targets to investigate and develop. One solution to this problem is to use predictive analysis techniques that can accurately identify cancer gene candidates in a timely fashion. Furthermore, the effort in identifying characteristics that are linked to cancer genes is crucial to further our understanding of this disease. These characteristics can be employed in recognising therapeutic drug targets. Here, we investigated whether certain genes properties can indicate the likelihood of it to be involved in the initiation or progression of cancer. We found that for cancer, the essentiality scores tend to be higher for cancer genes than for all protein coding human genes. A machine-learning model was developed and we found that essentiality related properties and properties arising from protein-protein interaction networks or evolution are particularly effective in predicting cancer-associated genes. We were also able to identify potential drug targets that have not been previously linked with cancer, but have the characteristics of cancer-related genes. Author SummaryMutations in numerous genes are known to be involved in cancer, yet there are undoubtedly many more to be discovered. We analysed a set of hundreds of cancer genes with the aim of finding out what makes them different from genes not known to be mutated in cancer. In particular, we found that genes that are essential for the survival of an organism are more likely to be involved in cancer. We used the gene properties that we examined to develop an artificial intelligence method that can accurately predict whether a gene is involved in cancer or not. Applying the method gives hundreds of non-cancer genes that resemble cancer genes. New discoveries of cancer genes are likely to be found within this set.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Novel ratio-metric features enable the identification of new driver genes across cancer types 97%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 96%
- DeepInsight-3D for precision oncology: an improved anti-cancer drug response prediction from high-dimensional multi-omics data with convolutional neural networks 95%
Similar papers in this journal
- Ranking Cancer Drivers via Betweenness-based Outlier Detection and Random Walks 97%
- scEvoNet: a gradient boosting-based method for prediction of cell state evolution 96%
- pyCancerSig: subclassifying human cancer with comprehensive single nucleotide, structural and microsatellite mutational signature deconstruction from whole genome sequencing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.