Back

One particle per residue is sufficient to describe all-atom protein structures

Heo, L.; Feig, M.

2023-05-23 biophysics
10.1101/2023.05.22.541652 bioRxiv
Show abstract

Atomistic resolution is considered the standard for high-resolution biomolecular structures, but coarse-grained models are often necessary to reflect limited experimental resolution or to achieve feasibility in computational studies. It is generally assumed that reduced representations involve a loss of detail, accuracy, and transferability. This study explores the use of advanced machine-learning networks to learn from known structures of proteins how to reconstruct atomistic models from reduced representations to assess how much information is lost when the vast knowledge about protein structures is taken into account. The main finding is that highly accurate and stereochemically realistic all-atom structures can be recovered with minimal loss of information from just a single bead per amino acid residue, especially when placed at the side chain center of mass. High-accuracy reconstructions with better than 1 [A] heavy atom root-mean square deviations are still possible when only C coordinates are used as input. This suggests that lower-resolution representations are essentially sufficient to represent protein structures when combined with a machine-learning framework that encodes knowledge from known structures. Practical applications of this high-accuracy reconstruction scheme are illustrated for adding atomistic detail to low-resolution structures from experiment or coarse-grained models generated from computational modeling. Moreover, a rapid, deterministic all-atom reconstruction scheme allows the implementation of an efficient multi-scale framework. As a demonstration, the rapid refinement of accurate models against cryoEM densities is shown where sampling at the coarse-grained level is guided by map correlation functions applied at the atomistic level. With this approach, the accuracy of standard all-atom simulation based refinement schemes can be matched at a fraction of the computational cost. STATEMENT OF SIGNIFICANCEThe fundamental insight of this work is that atomistic detail of proteins can be recovered with minimal loss of information from highly reduced representations with just a single bead per amino acid residue. This is possible by encoding the existing knowledge about protein structures in a machine-learning model. This suggests that it is not strictly necessary to resolve structures in atomistic detail in experiments, computational modeling, or the generation of protein conformations via neural networks since atomistic details can inferred quickly via the neural network. This increases the relevance of experimental structures obtained at lower resolutions and broadens the impact of coarse-grained modeling.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Acta Crystallographica Section D Structural Biology
54 papers in training set
Top 0.1%
22.7%
2
PLOS Computational Biology
1633 papers in training set
Top 2%
14.8%
3
IUCrJ
29 papers in training set
Top 0.1%
10.2%
4
Journal of Structural Biology
58 papers in training set
Top 0.2%
7.2%
50% of probability mass above
5
Structure
175 papers in training set
Top 0.6%
4.3%
6
Journal of Structural Biology: X
15 papers in training set
Top 0.1%
4.3%
7
Scientific Reports
3102 papers in training set
Top 36%
3.6%
8
Frontiers in Molecular Biosciences
100 papers in training set
Top 1%
1.8%
9
Journal of Applied Crystallography
14 papers in training set
Top 0.1%
1.8%
10
eLife
5422 papers in training set
Top 41%
1.7%
11
PLOS ONE
4510 papers in training set
Top 54%
1.7%
12
Biophysical Journal
545 papers in training set
Top 3%
1.7%
13
The Journal of Physical Chemistry B
158 papers in training set
Top 1%
1.3%
14
Communications Biology
886 papers in training set
Top 14%
1.2%
15
Journal of Chemical Information and Modeling
207 papers in training set
Top 2%
1.1%
16
Physical Biology
43 papers in training set
Top 2%
1.0%
17
Biology Methods and Protocols
53 papers in training set
Top 2%
1.0%
18
Bioinformatics
1061 papers in training set
Top 9%
0.9%
19
Computational and Structural Biotechnology Journal
216 papers in training set
Top 9%
0.8%
20
Nature Communications
4913 papers in training set
Top 64%
0.7%
21
Proteins: Structure, Function, and Bioinformatics
82 papers in training set
Top 1%
0.7%
22
Biophysical Reports
36 papers in training set
Top 0.5%
0.7%
23
International Journal of Molecular Sciences
453 papers in training set
Top 18%
0.6%
24
Proceedings of the National Academy of Sciences
2130 papers in training set
Top 47%
0.6%
25
SoftwareX
15 papers in training set
Top 0.5%
0.6%
26
NeuroImage
813 papers in training set
Top 7%
0.5%