Back

Hotspot coevolution at protein-protein interfaces is a key identifier of native protein complexes

Mishra, S.; Cooper, S. J.; Parks, J. M.; Mitchell, J. C.

2019-07-10 bioinformatics
10.1101/698233 bioRxiv
Show abstract

Protein-protein interactions play a key role in mediating numerous biological functions, with more than half the proteins in living organisms existing as either homo- or hetero-oligomeric assemblies. Protein subunits that form oligomers minimize the free energy of the complex, but exhaustive computational search-based docking methods have not comprehensively addressed the protein docking challenge of distinguishing a natively bound complex from non-native forms. In this study, we propose a scoring function, KFC-E, that accounts for both conservation and coevolution of putative binding hotspot residues at protein-protein interfaces. For a benchmark set of 53 bound complexes, KFC-E identifies a near-native binding mode as the top-scoring pose in 38% and in the top 5 in 55% of the complexes. For a set of 17 unbound complexes, KFC-E identifies a near-native pose in the top 10 ranked poses in more than 50% of the cases. By contrast, a scoring function that incorporates information on coevolution at predicted non-hotspots performs poorly by comparison. Our study highlights the importance of coevolution at hotspot residues in forming natively bound complexes and suggests a novel approach for coevolutionary scoring in protein docking.\n\nAuthor SummaryA fundamental problem in biology is to distinguish between the native and non-native bound forms of protein-protein complexes. Experimental methods are often used to detect the native bound forms of proteins but, are demanding in terms of time and resources. Computational approaches have proven to be a useful alternative; they sample the different binding configurations for a pair of interacting proteins and then use an heuristic or physical model to score them. In this study we propose a new scoring approach, KFC-E, which focuses on the evolutionary contributions from a subset of key interface residues (hotspots) to identify native bound complexes. KFC-E capitalizes on the wealth of information in protein sequence databases by incorporating residue-level conservation and coevolution of putative binding hotspots. As hotspot residues mediate the binding energetics of protein-protein interactions, we hypothesize that the knowledge of putative hotspots coupled with their evolutionary information should be helpful in the identification of native bound protein-protein complexes.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.1%
26.4%
2
PLOS Computational Biology
1863 papers in training set
Top 1.0%
21.8%
3
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.8%
6.7%
50% of probability mass above
4
Bioinformatics
1204 papers in training set
Top 4%
6.2%
5
Protein Science
246 papers in training set
Top 0.7%
5.5%
6
Journal of Molecular Biology
232 papers in training set
Top 0.6%
4.0%
7
Bioinformatics Advances
203 papers in training set
Top 2%
3.2%
8
Nature Communications
5641 papers in training set
Top 42%
2.1%
9
PLOS ONE
5266 papers in training set
Top 45%
2.1%
10
Scientific Reports
3612 papers in training set
Top 54%
1.7%
11
eLife
5828 papers in training set
Top 52%
1.5%
12
Molecular Biology and Evolution
542 papers in training set
Top 4%
1.5%
13
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
14
Biophysical Journal
631 papers in training set
Top 4%
1.1%
15
Communications Biology
993 papers in training set
Top 25%
1.0%
16
The Journal of Physical Chemistry B
167 papers in training set
Top 2%
0.9%
17
Journal of Chemical Theory and Computation
140 papers in training set
Top 1%
0.8%
18
ACS Omega
105 papers in training set
Top 4%
0.8%
19
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.8%
20
Protein Engineering, Design and Selection
15 papers in training set
Top 0.2%
0.8%
21
Journal of Cheminformatics
29 papers in training set
Top 0.8%
0.6%
22
Frontiers in Immunology
638 papers in training set
Top 11%
0.6%
23
Structure
193 papers in training set
Top 3%
0.6%