Back

The Gompertz curve for estimating growth rates of Protein Data Bank and protein folds

Sato, K.; TOMII, K.

2026-06-26 bioinformatics
10.64898/2026.06.24.732253 bioRxiv
Show abstract

The Protein Data Bank (PDB) is an ever-growing, open-access repository of structural data of biological molecules. This international database has been instrumental in the development of artificial intelligence and deep learning models for protein structure prediction and design. The PDB growth is a crucially important factor influencing further development of these models. Therefore, after analyzing the growth trend in PDB depositions since the archive's launch, we found that it is well fitted by the Gompertz function, a growth curve used across various disciplines. Furthermore, we observed that the function captures the "discovery of novel folds", i.e., the cumulative number of distinct folds among protein domains that constitute most of the PDB. Consequently, based on the fitting results, we estimated the likely numbers of PDB entries and protein folds. These findings provide insights into deceleration of growth in recent years and enable us to assess anticipated trends.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.2%
18.6%
2
Proteins: Structure, Function, and Bioinformatics
88 papers in training set
Top 0.1%
6.8%
3
Scientific Reports
3612 papers in training set
Top 10%
6.8%
4
Bioinformatics
1204 papers in training set
Top 4%
5.6%
5
BMC Bioinformatics
457 papers in training set
Top 2%
4.1%
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 1%
4.1%
7
PLOS ONE
5266 papers in training set
Top 36%
3.5%
8
Briefings in Bioinformatics
354 papers in training set
Top 3%
3.2%
50% of probability mass above
9
PLOS Computational Biology
1863 papers in training set
Top 11%
2.8%
10
Communications Biology
993 papers in training set
Top 9%
2.4%
11
Frontiers in Bioinformatics
49 papers in training set
Top 0.3%
2.1%
12
Molecules
39 papers in training set
Top 0.5%
2.0%
13
PeerJ
308 papers in training set
Top 5%
1.9%
14
The Journal of Physical Chemistry B
167 papers in training set
Top 1%
1.7%
15
Protein Science
246 papers in training set
Top 2%
1.7%
16
Journal of Molecular Biology
232 papers in training set
Top 2%
1.7%
17
GigaScience
212 papers in training set
Top 2%
1.7%
18
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.7%
19
Journal of Cheminformatics
29 papers in training set
Top 0.5%
1.3%
20
Scientific Data
209 papers in training set
Top 2%
1.1%
21
ACS Omega
105 papers in training set
Top 2%
1.1%
22
Journal of Computational Chemistry
13 papers in training set
Top 0.2%
1.1%
23
Database
61 papers in training set
Top 0.6%
1.1%
24
Computers in Biology and Medicine
128 papers in training set
Top 3%
1.1%
25
Computational Biology and Chemistry
28 papers in training set
Top 0.7%
1.1%
26
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
27
BioData Mining
22 papers in training set
Top 0.7%
0.9%
28
Journal of Computational Biology
48 papers in training set
Top 1.0%
0.9%
29
International Journal of Molecular Sciences
494 papers in training set
Top 15%
0.9%
30
International Journal of Biological Macromolecules
76 papers in training set
Top 2%
0.9%