Back

GeomMotif: A Benchmark for Arbitrary Geometric Preservation in Protein Generation

Strashnov, P.; Shevtsov, A.; Meshchaninov, V.; Kardymon, O.; Vetrov, D.

2026-04-22 bioinformatics
10.64898/2026.04.18.719378 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWMotif scaffolding in protein design involves generating complete protein structures while preserving the 3D geometry of designated structural fragments, analogous to image outpainting in computer vision. Current benchmarks focus on functional motifs, leaving general geometric preservation capabilities largely untested. We introduce GeomMotif, a systematic benchmark that evaluates arbitrary structural fragment preservation without requiring functional specificity. We construct 57 benchmark tasks, each containing one or two motifs with up to 7 continuous fragments, by sampling from the Protein Data Bank (PDB) to ensure a ground-truth, solvable conformation for every problem. The tasks are characterized by comprehensive structural and physicochemical properties: size, geometric context, secondary structure, hydrophobicity, charge, and degree of burial. These features enable detailed performance analysis beyond simple success rates, revealing model-specific strengths and limitations. We evaluate models using scRMSD and pLDDT for geometric fidelity and clustering for structural diversity and novelty. Our results show that sequence-based and structure-based approaches find different tasks challenging, and that geometric preservation varies significantly with structural and physicochemical context. GeomMotif provides insights complementary to function-focused benchmarks and establishes a foundation for improving protein generative models.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

1
Cell Systems
167 papers in training set
Top 1%
10.0%
2
Bioinformatics
1061 papers in training set
Top 3%
8.4%
3
Nature Communications
4913 papers in training set
Top 23%
8.3%
4
PLOS Computational Biology
1633 papers in training set
Top 5%
6.8%
5
Nature Methods
336 papers in training set
Top 2%
6.8%
6
Proceedings of the National Academy of Sciences
2130 papers in training set
Top 14%
4.8%
7
Journal of Cheminformatics
25 papers in training set
Top 0.2%
3.6%
8
Nature Machine Intelligence
61 papers in training set
Top 1.0%
3.6%
50% of probability mass above
9
Journal of Chemical Information and Modeling
207 papers in training set
Top 1%
3.2%
10
Briefings in Bioinformatics
326 papers in training set
Top 2%
3.0%
11
Bioinformatics Advances
184 papers in training set
Top 2%
2.6%
12
Nature Biotechnology
147 papers in training set
Top 4%
2.3%
13
Nucleic Acids Research
1128 papers in training set
Top 9%
2.1%
14
Protein Science
221 papers in training set
Top 0.7%
2.1%
15
Scientific Reports
3102 papers in training set
Top 50%
2.1%
16
Nature Computational Science
50 papers in training set
Top 0.5%
1.9%
17
BMC Bioinformatics
383 papers in training set
Top 5%
1.5%
18
Nature
575 papers in training set
Top 12%
1.5%
19
Computational and Structural Biotechnology Journal
216 papers in training set
Top 6%
1.3%
20
Journal of Structural Biology
58 papers in training set
Top 1.0%
1.3%
21
PLOS ONE
4510 papers in training set
Top 60%
1.2%
22
Science
429 papers in training set
Top 17%
1.2%
23
Proteins: Structure, Function, and Bioinformatics
82 papers in training set
Top 0.7%
0.9%
24
Advanced Science
249 papers in training set
Top 17%
0.9%
25
Cell Reports Methods
141 papers in training set
Top 5%
0.8%
26
Nature Genetics
240 papers in training set
Top 8%
0.7%
27
iScience
1063 papers in training set
Top 35%
0.7%
28
Protein Engineering, Design and Selection
14 papers in training set
Top 0.1%
0.6%
29
Communications Biology
886 papers in training set
Top 29%
0.6%
30
Genome Biology
555 papers in training set
Top 9%
0.6%