protPheMut: An Interpretable Machine Learning Tool for Classification of Cancer and Neurodevelopmental Disorders in Human Missense Variants
Wang, J.; Yang, M.; Zong, C.; Verkhivker, G.; Xiao, F.; Hu, G.
Show abstract
MotivationRecent advances in human genomics have revealed that missense mutations in a single protein can lead to distinctly different phenotypes. In particular, some mutations in oncoproteins like Ras, MEK, PI3K, PTEN, and SHP2 are linked various cancers and Neurodevelopmental Disorders (NDDs). While numerous tools exist for predicting the pathogenicity of missense mutations, linking these variants to certain phenotypes remains a major challenge, particularly in the context of personalized medicine. ResultsTo fill this gap, we developed protPheMut (Protein Phenotypic Mutations Analyzer), leveraging multiple interpretable machine learning methods and integrate diverse biophysics and network dynamics-based signatures, for the prediction of mutations of the same protein can promote cancer, or NDDs. We illustrate the utility of protPheMut in phenotypes (cancer/NDDs) prediction by the mutation analysis of two protein cases, that are PI3K and PTEN. Compared to seven other predictive tools, protPheMut demonstrated exceptional accuracy in forecasting phenotypic effects, achieving an AUROC of 0.8501 for PI3K mutations related to cancer and Cowden syndrome. For multi-phenotypes prediction of PTEN mutations related to cancer, PHTS, and HCPS, protPheMut achieved an AUC of 0.9349 through micro-averaging. Using SHAP model explanations, we gained insights into the mechanisms driving phenotype formation. A userfriendly website deployment is also provided. AvailabilitySource code and data are available at https://github.com/Spencer-JRWang/protPheMut. We also provide a user-friendly website at http://netprotlab.com/protPheMut. Supplementary informationSupplementary data are available at Bioinformatics online. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=64 SRC="FIGDIR/small/631365v1_ufig1.gif" ALT="Figure 1"> View larger version (28K): org.highwire.dtl.DTLVardef@df55e3org.highwire.dtl.DTLVardef@7fe3f9org.highwire.dtl.DTLVardef@501cccorg.highwire.dtl.DTLVardef@192c98e_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rhapsody: Pathogenicity prediction of human missense variants based on protein sequence, structure and dynamics 95%
- 3Cnet: Pathogenicity prediction of human variants using knowledge transfer with deep recurrent neural networks 94%
- MutaFrame - an interpretative visualization framework for deleteriousness prediction of missense variants in the human exome 94%
Similar papers in this journal
Similar papers in this journal
- Cancer SIGVAR: A semi-automated interpretation tool for germline variants of hereditary cancer-related genes 94%
- ModelMatcher: A scientist-centric online platform to facilitate collaborations between stakeholders of rare and undiagnosed disease research 92%
- Matching whole genomes to rare genetic disorders: Identification of potential causative variants using phenotype-weighted knowledge in the CAGI SickKids5 clinical genomes challenge 92%
Similar papers in this journal
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 94%
- Improved model quality assessment using sequence and structural information by enhanced deep neural networks 93%
- AlphaFold2-aware protein-DNA binding site prediction using graph transformer 93%
Similar papers in this journal
- DisPhaseDB, an integrative database of diseases related variations in liquid-liquid phase separation proteins 94%
- Getting to know each other: PPIMem, a novel approach for predicting transmembrane protein-protein complexes 94%
- A Multi-Layered Computational Structural Genomics Approach Enhances Domain-Specific Interpretation of Kleefstra Syndrome Variants in EHMT1 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.