Uni-Fold MuSSe: De Novo Protein Complex Prediction with Protein Language Models
Zhu, J.; He, Z.; Li, Z.; Ke, G.; Zhang, L.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWAccurately solving the structures of protein complexes is crucial for understanding and further modifying biological activities. Recent success of AlphaFold and its variants shows that deep learning models are capable of accurately predicting protein complex structures, yet with the painstaking effort of homology search and pairing. To bypass this need, we present Uni-Fold MuSSe (Multimer with Single Sequence inputs), which predicts protein complex structures from their primary sequences with the aid of pre-trained protein language models. Specifically, we built protein complex prediction models based on the protein sequence representations of ESM-2, a large protein language model with 3 billion parameters. In order to adapt the language model to inter-protein evolutionary patterns, we slightly modified and further pre-trained the language model on groups of protein sequences with known interactions. Our results highlight the potential of protein language models for complex prediction and suggest room for improvements.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- ProtMamba: a homology-aware but alignment-free protein state space model 98%
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 97%
- QDeep: distance-based protein model quality estimation by residue-level ensemble error classifications using stacked deep residual neural networks 97%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.