ProteinAligner: A Multi-modal Pretraining Framework for Protein Foundation Models
Zhang, L.; Guo, H.; Schaffer, L. V.; Ko, Y. S.; Singh, D.; Rahmani, H.; Grotjahn, D.; Villa, E.; Gilson, M.; Wang, W.; Ideker, T.; Xing, E.; Xie, P.
Show abstract
Protein foundation models, particularly protein language models, have demonstrated strong success in learning meaningful representations of proteins using transformer architectures pretrained on large-scale protein datasets with self-supervised learning. These representations have been highly effective for downstream tasks such as predicting protein functions and properties. However, most current protein foundation models focus on pretraining with amino acid sequences, often neglecting additional modalities like protein structures and related literature, both of which provide valuable insights. To address this gap, we propose a multi-modal pretraining approach that integrates three key modalities - protein sequences, structures, and literature text. In our framework, the protein sequence modality serves as the anchor, with the other two modalities aligned to it, enhancing the models capacity to capture more comprehensive protein information. ProteinAligner out-performed state-of-the-art protein foundation models in predicting protein functions and properties across diverse down-stream tasks.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accelerating protein engineering with fitness landscape modeling and reinforcement learning 96%
- Generalized Biological Foundation Model with Unified Nucleic Acid and Protein Language 96%
- PepFlow: direct conformational sampling from peptide energy landscapes through hypernetwork-conditioned diffusion 94%
Similar papers in this journal
- INTREPPPID - An Orthologue-Informed Quintuplet Network for Cross-Species Prediction of Protein-Protein Interaction 96%
- Protein engineering via Bayesian optimization-guided evolutionary algorithm and robotic experiments 95%
- GraphCPLMQA: Assessing protein model quality based on deep graph coupled networks using protein language model 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.