Task-Specialized Protein Language Models Decode the Sequence Grammar of Post-Translational Modification Sites
Adhikari, S.; Mondal, J.
Show abstract
Post-translational modifications (PTMs) regulate protein signaling, localization, degradation, and cellular decision-making, yet the sequence determinants that distinguish modified from chemically eligible but unmodified residues remain difficult to decode at proteome scale. Here, we examine whether adapting a general protein language model to PTM-site prediction can reveal the biochemical logic underlying residue-level modification. We fine-tune ESM2, a protein language model trained on tens of millions of evolutionarily diverse protein sequences, for phosphorylation, acetylation, and ubiquitination-site prediction. To address the pronounced class imbalance inherent in proteome-wide PTM annotation, we combine parameter-efficient fine-tuning with focal-loss training. The resulting task-specialized models show that PTM recognition depends on model capacity, annotation depth, and modification chemistry: phosphorylation benefits from larger models, whereas acetylation and ubiquitination peak at intermediate scale. More importantly, the fine-tuned phosphorylation model exposes three layers of biological organization: it recovers canonical kinase-recognition motifs without kinase-label supervision, resolves pathway-level functional relationships among proteins from sequence-derived embeddings, and preserves evolutionary signatures of homologous phosphorylation sites across 200 eukaryotic species. These results establish task-specialized protein language models as interpretable instruments for probing PTM-site biochemistry, kinase specificity, functional organization, and evolutionary conservation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Structure-Based Function Prediction using Graph Convolutional Networks 96%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
- Generalizable and scalable protein stability prediction with rewired protein generative models 96%
Similar papers in this journal
- Direct prediction of intrinsically disordered protein conformational properties from sequence 96%
- Sliding Window INteraction Grammar (SWING): a generalized interaction language model for peptide and protein interactions 96%
- Self-Supervised Deep-Learning Encodes High-Resolution Features of Protein Subcellular Localization 96%
Similar papers in this journal
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 94%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 94%
- Discovering molecular features of intrinsically disordered regions by using evolution for contrastive learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.