Extending Prot2Token: Aligning Protein Language Models for Unified and Diverse Protein Prediction Tasks
Pourmirzaei, M.; Han, Y.; Esmaili, F.; Pourmirzaei, M.; Alqarghuli, S.; Chen, K.; Xu, D.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWComprehensive protein function and property prediction remains a major challenge due to the vast diversity of sequences, structural variations, and limited labeled data. Existing models are often specialized to be task-specific, requiring independent training, which limits scalability. To address this, we extend Prot2Token, a unified autoregressive framework that focuses on the post-training alignment of pre-trained protein language models (PLMs), to new applications. Our approach enables next-token prediction across new applications of proteinprediction tasks, including protein-protein structure similarity, 3D structure prediction, mutation stability, post-translational modifications (PTMs), substratekinase phosphorylation sites, protein-protein affinity, and protein-ion binding sites. We introduce a self-supervised pre-training stage for the decoder, enhancing model initialization and improving downstream predictions. By integrating a causal autoregressive transformer with a pre-trained ESM-2 encoder, our model effectively aligns diverse protein tasks within a single framework. Additionally, we discuss the opportunities and limitations of this approach, providing insights for future research in optimizing PLMs as a general tool for broader biological applications. Code is available on GitHub Repository.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- All-Atom Protein Sequence Design using Discrete Diffusion Models 95%
- Sequence-based Drug-Target Complex Pre-training Enhances Protein-Ligand Binding Process Predictions Tackling Crypticity 94%
- Structure-aware Protein Solubility Prediction From Sequence Through Graph Convolutional Network And Predicted Contact Map 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.