Back

Protein Structure Description with rho, theta and phi: A Case Study with Caenopore-5

Li, W.

2025-01-10 bioinformatics
10.1101/2025.01.08.631847 bioRxiv
Show abstract

Since its establishment in 1971, the Protein Data Bank (PDB) has been using Cartesian coordinate system (CCS) as the standard framework for protein structure description with x, y, z. Despite the interconvertibility of CCS and spherical coordinate systems (SCS,{rho} , {theta} and{phi} ), CCS remains to date the default and the only framework for protein structure description in PDB. Recent advances in protein structure prediction (e.g., AlphaFold) revolutionized the field by integrating deep learning algorithms with experimental structural data, achieving unprecedented accuracy of protein structure prediction and relying on Cartesian representation of protein structures to extract geometric features. To this end, questions remain about what drives the next stage of continued performance improvement of protein structure prediction. Therefore, this article introduces an alternative coordinate system for protein structure description and feature extraction. Using Caenopore-5 as an example, this article redefines protein backbone structures using atomic bonding networks (ABN) within the SCS framework (ABN-SCS), leading to the extraction of a set of spherical parameters ({rho}, {theta} and{phi} ) from the NMR ensemble of Caenopore-5, encompassing 477 covalent bonds and 80 peptide bonds within its backbone for each structural model in its NMR ensemble. Finally, this work demonstrates that ABN-SCS enables characterization of spherical bond-level geometries, expanding the feature space available for computational pipelines such as AlphaFold2, and argues that integrating ABN-SCS features into protein structure prediction pipelines can enhance geometric fidelity, and that the time is now ripe for the trapped spherical features [{rho}, {theta}, {phi}] to be integrated into algorithms such as AF2 towards protein structure prediction with improved performance.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Molecules
39 papers in training set
Top 0.1%
15.3%
2
Briefings in Bioinformatics
354 papers in training set
Top 0.6%
9.9%
3
International Journal of Molecular Sciences
494 papers in training set
Top 0.7%
6.8%
4
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.7%
6.8%
5
PLOS ONE
5266 papers in training set
Top 28%
5.6%
6
Journal of Molecular Biology
232 papers in training set
Top 0.4%
4.9%
7
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.8%
4.4%
50% of probability mass above
8
PLOS Computational Biology
1863 papers in training set
Top 10%
3.3%
9
Biochimica et Biophysica Acta (BBA) - Biomembranes
36 papers in training set
Top 0.1%
3.3%
10
Biomolecules
100 papers in training set
Top 0.4%
2.8%
11
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.4%
2.8%
12
Scientific Reports
3612 papers in training set
Top 39%
2.7%
13
Computers in Biology and Medicine
128 papers in training set
Top 2%
2.4%
14
Journal of Molecular Graphics and Modelling
17 papers in training set
Top 0.2%
2.2%
15
Journal of Computational Chemistry
13 papers in training set
Top 0.1%
1.5%
16
International Journal of Biological Macromolecules
76 papers in training set
Top 1%
1.1%
17
Frontiers in Molecular Biosciences
102 papers in training set
Top 1%
1.1%
18
Computational Biology and Chemistry
28 papers in training set
Top 0.8%
1.1%
19
Bioinformatics
1204 papers in training set
Top 9%
0.9%
20
Journal of Chemical Theory and Computation
140 papers in training set
Top 1%
0.9%
21
The Journal of Physical Chemistry Letters
63 papers in training set
Top 0.7%
0.9%
22
Biochemical and Biophysical Research Communications
84 papers in training set
Top 3%
0.6%
23
Biophysics and Physicobiology
11 papers in training set
Top 0.1%
0.6%
24
Biochimie
25 papers in training set
Top 0.9%
0.6%
25
ACS Omega
105 papers in training set
Top 4%
0.6%