Protein Complex Structure Prediction Powered by Multiple Sequence Alignment of Interologs from Multiple Taxonomic Ranks and AlphaFold2
Si, Y.; Yan, C.
Show abstract
AlphaFold2 is expected to be able to predict protein complex structures as long as a multiple sequence alignment (MSA) of the interologs of the target protein-protein interaction (PPI) can be provided. In this study, a simplified phylogeny-based approach was applied to generate the MSA of interologs, which was then used as the input to AlphaFold2 for protein complex structure prediction. Extensively benchmarked this protocol on non-redundant PPI dataset including 107 bacterial PPIs and 442 eukaryotic PPIs, we show complex structures of 79.5% of the bacterial PPIs and 49.8% of the eukaryotic PPIs can be successfully predicted, which yielded significantly better performance than the application of MSA of interologs prepared by two existing approaches. Considering PPIs may not be conserved in species with long evolutionary distances, we further restricted interologs in the MSA to different taxonomic ranks of the species of the target PPI in protein complex structure prediction. We found the success rates can be increased to 87.9% for the bacterial PPIs and 56.3% for the eukaryotic PPIs if interologs in the MSA are restricted to a specific taxonomic rank of the species of each target PPI. Finally, we show the optimal taxonomic ranks for protein complex structure prediction can be selected with the application of the predicted TM-scores of the output models.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GraphGPSM: a global scoring model for protein structure using graph neural networks 98%
- Improved model quality assessment using sequence and structural information by enhanced deep neural networks 97%
- Improved inter-protein contact prediction using dimensional hybrid residual networks and protein language models 97%
Similar papers in this journal
- Pathfinder: protein folding pathway prediction based on conformational sampling 96%
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 96%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 95%
Similar papers in this journal
- Structural analogue-based protein structure domain assembly assisted by deep learning 97%
- DeepUMQA: Ultrafast Shape Recognition-based Protein Model Quality Assessment using Deep Learning 97%
- A de novo protein structure prediction by iterative partition sampling, topology adjustment, and residue-level distance deviation optimization 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.