AstraPTM: Context-Aware PTM Prediction Model for Large-Scale Proteins
Bozkurt, C.; Goteti, A.
Show abstract
Post-translational modifications (PTMs) are critical molecular events that pro-foundly influence protein stability, localization, and function. While numerous computational tools exist for PTM site prediction, most struggle with handling large proteins and rely on separate models for each modification. To address these challenges, we introduce AstraPTM, a transformer-based framework that predicts 25 PTMs in a single pass. By leveraging advanced protein embeddings (ESM2) and training on a high-coverage dbPTM dataset, AstraPTM captures both short-range sequence motifs and long-range interactions across proteins without a sequence length limitation. AstraPTM combines a binary classification module--indicating whether a residue is modified--with a multi-label module that pinpoints specific PTM types. This dual approach achieves high accuracy on well-represented PTMs (e.g., phospho-rylation, glycosylation) while maintaining sensitivity for rarer modifications. In benchmarks against existing methods such as MusiteDeep and MIND-S, AstraPTM demonstrates competitive or superior performance, demonstrating AUC-ROC above 99% for well-represented modifications, underscoring its versatility for proteome-wide annotation. Beyond prediction, the models capacity to handle full-length proteins offers a powerful resource for researchers investigating PTM crosstalk and disease pathways, ultimately bridging the gap between large-scale omics data and targeted biomedical applications.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 97%
- ProteinBERT: A universal deep-learning model of protein sequence and function 96%
- DistilProtBert: A distilled protein language model used to distinguish between real proteins and their randomly shuffled counterparts 96%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 94%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 96%
- Designing diverse and high-performance proteins with a large language model in the loop 94%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.