ZIPPI: proteome-scale sequence-based evaluation of protein-protein interaction models
Zhao, H.; Petrey, D. S.; Murray, D.; Honig, B.
Show abstract
Predicting protein-protein interactions (PPI) is a challenging problem of central importance in fundamental biology. With the increasing number of available PPI prediction methods and databases, an effective evaluation model would be extremely valuable. Here we introduce ZIPPI (Z-score for Information about Protein-Protein Interfaces), which evaluates structural models of a complex based on sequence co-evolution and conservation involving residues that are in contact in the interface. The interface Z-score (ZIPPI score) is calculated by comparing metrics for interface contacts to metrics obtained from randomly chosen surface residues. Since contacting residues are defined by the structural model, this obviates the need of accounting for indirect interactions with methods such as Direct Coupling Analysis. Although ZIPPI relies on species-paired multiple sequence alignments, its focus on contacting interfacial residues and the avoidance of direct coupling methods makes it computationally efficient. The performance of ZIPPI is evaluated through applications to experimentally determined complexes from the Protein Data Bank (PDB) and to decoys from the Critical Assessment of PRedicted Interactions (CAPRI) experiment. We demonstrate how ZIPPI can be implemented on a genome-wide scale by calculating scores for millions of structural models of protein-protein interactions in the E. coli interactome as predicted by PrePPI. Many PrePPI predictions filtered by ZIPPI score are novel. In all, this proteome-scale method shows promising feasibility for applications to the full human protein interactome, which is not yet accessible to deep learning methods.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Knot or Not? Sequence-Based Identification of Knotted Proteins With Machine Learning 96%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 96%
- RosettaDDGPrediction for high-throughput mutational scans: from stability to binding 96%
Similar papers in this journal
- PIPENN-EMB: ensemble net and protein embeddings generalise protein interface prediction beyond homology 97%
- ProteinGLUE: A multi-task benchmark suite for self-supervised protein modeling. 97%
- Evaluating the Significance of Embedding-Based Protein Sequence Alignment with Clustering and Double Dynamic Programming for Remote Homology 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.