Accurate Prediction Of Asparagine Deamidation In Biologics Using Advanced Machine Learning Models
Ahmed, S.; Swope, N.; Stanev, V.; Hofele, R.; WuDunn, D.; Jain, R.; Delmar, J.; Pouryahya, M.
Show abstract
The spontaneous deamidation of asparagine residues remains a major obstacle to the stability and efficacy of protein therapeutics. Currently available models in the literature for predicting deamidation liabilities can suffer from limited generalizability, likely due to biases such as sequence similarity within datasets. In this study, we built machine learning models using protein language models (e.g., ESM2) and graph neural networks (GNNs), trained on a comprehensive dataset of 591 asparagine sites from over 105 protein molecules. To address the critical issue of data leakage, we implemented a peptide grouping strategy yielding more accurate estimates of model performance for novel deamidation sites. Our analysis shows that, when sequence similarity bias is controlled, protein language models match traditional feature-based models that use amino acid composition, k-mers, PSSMs, and predicted secondary structure/solvent accessibility, while offering substantial computational advantages. Additionally, our GNN-based pipeline further increases prediction accuracy by up to 8% compared to language model-only tools and delivers a 15-25% improvement over motif-based approaches. This methodological framework enables more reliable and rapid in-silico prediction of deamidation liabilities, potentially reducing costly late-stage interventions in protein therapeutic development and is generalizable to the modeling of additional protein post-translational modifications.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- BAGEL: Protein Engineering via Exploration of an Energy Landscape 95%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 95%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 95%
Similar papers in this journal
Similar papers in this journal
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 97%
- Learning Context-aware Structural Representations to Predict Antigen and Antibody Binding Interfaces 96%
- Enhancing predictions of protein stability changes induced by single mutations using MSA-based Language Models 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.