Clade-wise alignment integration improves co-evolutionary signals for protein-protein interaction prediction
Fang, T.; Szklarczyk, D.; Hachilif, R.; von Mering, C.
Show abstract
BackgroundProtein-protein interactions play essential roles in almost all biological processes. The binding interfaces between interacting proteins impose evolutionary constraints, leading to co-evolutionary signals that have successfully been employed to predict protein interactions from multiple sequence alignments (MSAs). During the construction of MSAs for this purpose, critical choices have to be made: how to ensure the reliable identification of orthologs, how to deal with paralogs, and how to optimally balance the need for large alignments versus sufficient alignment quality. ResultsHere, we propose a divide-and-conquer strategy for MSA generation: instead of building a single, large alignment for each protein, multiple distinct alignments are constructed, each covering only a single clade in the tree of life. Co-evolutionary signals are searched separately within these clades, and are only subsequently integrated into a final interaction prediction using machine learning. We find that this strategy markedly improves overall prediction performance, concomitant with better alignment quality. Using the popular DCA algorithm to systematically search pairs of such alignments, a genome-wide all-against-all interaction scan in a bacterial genome is demonstrated. ConclusionsGiven the recent successes of AlphaFold in predicting protein-protein interactions at atomic detail, a discover-and-refine approach is proposed: our method could provide a fast and accurate strategy for pre-screening the entire genome, submitting to AlphaFold only promising interaction candidates - thus reducing false positives as well as computation time.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 95%
- Knowledge-guided data mining on the standardized architecture of NRPS: subtypes, novel motifs, and sequence entanglements 95%
- From complete cross-docking to partners identification and binding sites predictions 95%
Similar papers in this journal
- Cracking the black box of deep sequence-based protein-protein interaction prediction 97%
- PSTP: Decoding Latent Sequence Grammar for Protein Phase Separation through Transfer Learning and Attention 95%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 94%
Similar papers in this journal
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 96%
- PIFiA: Self-supervised Approach for Protein Functional Annotation from Single-Cell Imaging Data 94%
- hu.MAP 2.0: Integration of over 15,000 proteomic experiments builds a global compendium of human multiprotein assemblies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.