Improve the Protein Complex Prediction with Protein Language Models
Chen, B.; Xie, Z.; Qiu, J.; Ye, Z.; Xu, J.; Tang, J.
Show abstract
AlphaFold-Multimer has greatly improved protein complex structure prediction, but its accuracy also depends on the quality of the multiple sequence alignment (MSA) formed by the interacting homologs (i.e., interologs) of the complex under prediction. Here we propose a novel method, denoted as ESMPair, that can identify interologs of a complex by making use of protein language models (PLMs). We show that ESMPair can generate better interologs than the default MSA generation method in AlphaFold-Multimer. Our method results in better complex structure prediction than AlphaFold-Multimer by a large margin (+10.7% in terms of the Top-5 best DockQ), especially when the predicted complex structures have low confidence. We further show that by combining several MSA generation methods, we may yield even better complex structure prediction accuracy than Alphafold-Multimer (+22% in terms of the Top-5 best DockQ). We systematically analyze the impact factors of our algorithm and find out the diversity of MSA of interologs significantly affects the prediction accuracy. Moreover, we show that ESMPair performs particularly well on complexes in eucaryotes.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- High-accuracy protein complex structure modeling based on sequence-derived structure complementarity 97%
- Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction 97%
- Structure-Based Function Prediction using Graph Convolutional Networks 97%
Similar papers in this journal
- GraphCPLMQA: Assessing protein model quality based on deep graph coupled networks using protein language model 96%
- Predicting the structures of cyclic peptides containing unnatural amino acids by HighFold2 95%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 95%
Similar papers in this journal
- Full-length de novo protein structure determination from cryo-EM maps using deep learning 95%
- OPUS-X: An Open-Source Toolkit for Protein Torsion Angles, Secondary Structure, Solvent Accessibility, Contact Map Predictions, and 3D Folding 95%
- Inference of genome 3D architecture by modeling overdispersion of Hi-C data 94%
Similar papers in this journal
- Improving the prediction of protein stability changes upon mutations by geometric learning and a pre-training strategy 97%
- Fast and effective protein model refinement by deep graph neural networks 96%
- ECloudGen: Leveraging Electron Clouds as a Latent Variable to Scale Up Structure-based Molecular Design 92%
Similar papers in this journal
- Deep Template-based Protein Structure Prediction 96%
- Deducing high-accuracy protein contact-maps from a triplet of coevolutionary matrices through deep residual convolutional networks 95%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.