Integrative machine learning predicts activating kinase mutations for precision oncology
Wang, Y.; Wan, F.; Chen, Z.; Nukpexzah, J.; Pan, I. T.; Stebe, K. J.; de la Fuente-Nunez, C.; Radhakrishnan, R.
Show abstract
Kinases are enzymes that catalyze phosphorylation and play crucial roles in a myriad of cellular regulatory processes and hemostasis. Patient-specific genetic mutations that aberrantly activate kinases can profoundly influence cancer progression and alter drug efficacy. Predicting the impact of such missense mutations across the human kinome on protein function and cellular signaling is therefore a critical step toward personalized targeted therapy. Here, we present Kinome-AI, an integrative machine learning framework that classifies kinase missense mutations as activating or non-activating. Kinome-AI is trained on a rich multi-modal feature set, including residue-level biochemical changes, sequence embeddings from a protein language model, and structural descriptors of kinase-ATP-substrate complexes derived from molecular modeling. Notably, detailed structural features were available for only 21% of mutants; we leverage these as privileged information during training to impute missing structural data for the remaining [~]79. This strategy boosts performance without requiring structural inputs for new (unseen) mutations. The resulting classifier achieves an area under the receiver operating characteristic curve (AUROC) of 0.85 and a balanced accuracy (BACC) of 0.76 across 1,003 mutations spanning 110 different kinases --substantially outperforming existing bioinformatics and general-purpose variant effect predictors. This work provides a robust approach to quantify sequence-structure- function relationships of cancer-driving kinase mutations, paving the way for improved personalized cancer treatment. Significance StatementIn cancer patients, numerous mutations in diverse protein kinases lead to marked differences in disease progression and drug response. Identifying which kinase mutations are activating in individual patients is therefore critical for precision oncology. Drawing inspiration from teacher- student (privileged information) learning, we developed a deep learning framework that integrates structural features from molecular simulations with sequence embeddings from protein language models. This approach enables accurate binary classification of the activation status of kinase mutations. Our study demonstrates how data-driven algorithms can leverage accumulated sequence and structural knowledge of known mutations to predict the effects of novel variants a priori. The model, termed Kinome-AI, shows significant promise for incorporation into personalized cancer therapy decision pipelines.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 96%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 95%
- THLANet: A Deep Learning Framework for Predicting TCR-pHLA Binding in Immunotherapy Applications 95%
Similar papers in this journal
Similar papers in this journal
- PARROT: a flexible recurrent neural network framework for analysis of large protein datasets 94%
- Death by a Thousand Cuts -- Combining Kinase Inhibitors for Selective Target Inhibition and Rational Polypharmacology 94%
- Training deep neural density estimators to identify mechanistic models of neural dynamics 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.