Back

ProteinAligner: A Multi-modal Pretraining Framework for Protein Foundation Models

Zhang, L.; Guo, H.; Schaffer, L. V.; Ko, Y. S.; Singh, D.; Rahmani, H.; Grotjahn, D.; Villa, E.; Gilson, M.; Wang, W.; Ideker, T.; Xing, E.; Xie, P.

2024-10-06 bioinformatics
10.1101/2024.10.06.616870 bioRxiv
Show abstract

Protein foundation models, particularly protein language models, have demonstrated strong success in learning meaningful representations of proteins using transformer architectures pretrained on large-scale protein datasets with self-supervised learning. These representations have been highly effective for downstream tasks such as predicting protein functions and properties. However, most current protein foundation models focus on pretraining with amino acid sequences, often neglecting additional modalities like protein structures and related literature, both of which provide valuable insights. To address this gap, we propose a multi-modal pretraining approach that integrates three key modalities - protein sequences, structures, and literature text. In our framework, the protein sequence modality serves as the anchor, with the other two modalities aligned to it, enhancing the models capacity to capture more comprehensive protein information. ProteinAligner out-performed state-of-the-art protein foundation models in predicting protein functions and properties across diverse down-stream tasks.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.