Prediction of Kinase-Substrate Associations Using The Functional Landscape of Kinases and Phosphorylation Sites
Ayati, M.; Yilmaz, S.; Blasco Tavares Pereira Lopes, F.; Chance, M. R.; Koyuturk, M.
Show abstract
MotivationProtein phosphorylation is a key post-translational modification that plays a central role in many cellular processes. With recent advances in biotechnology, thousands of phosphorylated sites can be identified and quantified in a given sample, enabling proteome-wide screening of cellular signaling. However, the kinase(s) that phosphorylate most (> 90%) of the identified phosphorylation sites are unknown. Knowledge of kinase-substrate associations is also mostly limited to a small number of well-studied kinases, with 20% of known kinases accounting for the phosphorylation of 87% of currently annotated sites. The scarcity of available annotations calls for the development of computational algorithms for more comprehensive and reliable prediction of kinase-substrate associations. ResultsTo broadly utilize available structural, functional, evolutionary, and contextual information in predicting kinase-substrate associations, we develop a network-based machine learning framework. Our framework integrates a multitude of data sources to characterize the landscape of functional relationships and associations among phosphosites and kinases. To construct a phosphosite-phosphosite association network, we use sequence similarity, shared biological pathways, co-evolution, co-occurrence, and co-phosphorylation of phosphosites across different biological states. To construct a kinase-kinase association network, we integrate protein-protein interactions, shared biological pathways, and membership in common kinase families. We use node embeddings computed from these heterogeneous networks to train machine learning models for predicting kinase-substrate associations. Our systematic computational experiments using the PhosphositePLUS database shows that the resulting algorithm, NetKSA, outperforms state-of-the-art algorithms and resources, including KinomeXplorer and LinkPhinder, in reliably predicting KSAs. By stratifying the ranking of kinases, NetKSA also enables annotation of phosphosites that are targeted by relatively less-studied kinases. Finally, we observe that the performance of NetKSA is robust to the choice of network embedding algorithms, while each type of network contributes valuable information that is complementary to the information provided by other networks. ConclusionRepresentation of available functional information on kinases and phosphorylation sites, along with integrative machine learning algorithms, has the potential to significantly enhance our knowledge on kinase-substrate associations. AvailabilityThe code and data are available at compbio.case.edu/NetKSA.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- KSMoFinder - Knowledge graph embedding of proteins and motifs for predicting kinases of human phosphosites 98%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 94%
- ICoN: Integration using Co-attention across Biological Networks 93%
Similar papers in this journal
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 94%
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 93%
- CASTER-DTA: Equivariant Graph Neural Networks for Predicting Drug-Target Affinity 93%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 94%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 94%
- DAGBagM: Learning directed acyclic graphs of mixed variables with an application to identify prognostic protein biomarkers in ovarian cancer 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.