AstraPTM2: A Context-Aware Transformer for Broad-Spectrum PTM Prediction
Bozkurt, C.; Vasilyeva, A.; Goteti, A.
Show abstract
Post-translational modifications (PTMs) are covalent changes in proteins after biosynthesis that shape stability, localization, and function. While numerous computational tools exist for PTM site prediction, most struggle to handle full-length proteins without truncating context, focus on only a limited number of PTM types, and perform unevenly on rare modifications. We present AstraPTM2, a transformer-based model that predicts 39 distinct PTM types on full-length sequences. By combining ESM-2 embeddings, AlphaFold2-derived structural features, and protein-level descriptors, AstraPTM2 captures both short-range motifs and long-range dependencies. Training uses a three-stage curriculum and adaptive focal loss to balance rare and common PTMs, followed by per-label affine calibration and optimized thresholds for well-calibrated probabilities. In hold-out tests, AstraPTM2 achieves AUROC = 0.99 and macro-F1 = 59% across 39 PTM types, with particularly strong performance on rare motif-driven PTMs such as O-linked glycosylation and sumoylation. Results are available through the Orbion web platform, which offers synchronized 2D and 3D visualizations, dual prediction modes (calibrated and exploratory), and reproducible exports to support downstream experimental planning. AstraPTM2 can be accessed at https://www.orbion.life.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 97%
- ProteinBERT: A universal deep-learning model of protein sequence and function 97%
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 96%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.