Back

SNAC-DB: An ML-Ready Database for Antibody and NANOBODY(R) VHH-Antigen Complexes with Expanded Structural Diversity and Real-World Benchmarking

Gupta, A.; Munoz Rivero, B.; Li, R.; Roel-Touris, J.; Fomekong Nanfack, Y.; Wendt, M.; Qiu, Y.; Furtmann, N.

2026-04-26 bioinformatics
10.64898/2026.04.22.720253 bioRxiv
Show abstract

Predicting antibody and NANOBODY(R) VHH-antigen complexes remains a critical challenge for state-of-the-art structure prediction models, limiting their impact in therapeutic discovery pipelines. We introduce SNAC-DB, an ML-ready database and curation pipeline enriched with structural biology expertise, designed to accelerate model accuracy and generalization by providing 31-37% expanded structural diversity over existing resources like SAbDab through comprehensive re-curation that extracts maximum value from available experimental structures. SNAC-DB expands coverage by capturing often-overlooked complexes and accurately identifying complete multi-chain epitopes through improved biological-assembly-based logic. Built for ML practitioners, SNAC-DB provides standardized formats with multi-threshold structure-based clustering to enable principled sample weighting during training. Using a rigorous benchmark of public PDB entries deposited post-May 2024 plus confidential therapeutic structures, we evaluate seven leading models (Protenix-v1, OpenFold-3p2, RosettaFold-3, Boltz-2, Boltz-1x, Chai-1, and AlphaFold2.3-multimer) with evaluation methodology tailored to antibody/NAN-OBODY(R) VHH-antigen complexes to ensure correct handling of multi-chain epitopes, revealing systematic performance gaps: success rates rarely exceed 25%, confidence-based ranking fails to identify best predictions even when accurate structures exist in ensembles, and all models consistently struggle with therapeutically relevant NANOBODY(R) VHHs. Systematic evaluation of sampling strategies demonstrates that while generating 1000 samples per target substantially increases the likelihood of producing accurate structures (oracle selection improves from 11.9% to 50.5%), confidence-based ranking remains nearly flat (between 10.9% and 14.9%), revealing that improved ranking mechanisms represent a more tractable path to performance gains. Finally, fine-tuning GeoDock on SNAC-DB yields higher success rates than training on SAbDab (11.0% vs. 7.1% for antibodies; 7.0% vs. 4.0% for NANOBODY(R) VHHs), suggesting that SNAC-DBs expanded structural diversity translates to improved model generalization. Significance StatementComputational antibody/NANOBODY(R) VHH design shows promise but remains unreliable for therapeutic development. SNAC-DB provides 31-37% expanded structural diversity through comprehensive data curation, immediately accelerating model development. Benchmarking seven leading AI models reveals accuracy rarely exceeds 25% on therapeutic targets, with confidence-based ranking failing to identify correct structures even when they exist in model outputs. Training on SNAC-DB increases prediction accuracy, validating that high-quality, diverse training data is critical for advancing computational methods toward clinical impact.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
mAbs
32 papers in training set
Top 0.1%
30.8%
2
Bioinformatics
1204 papers in training set
Top 3%
7.8%
3
Protein Science
246 papers in training set
Top 0.5%
7.2%
4
Briefings in Bioinformatics
354 papers in training set
Top 1%
6.2%
50% of probability mass above
5
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.5%
5.5%
6
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
4.8%
7
PLOS Computational Biology
1863 papers in training set
Top 8%
4.3%
8
Bioinformatics Advances
203 papers in training set
Top 1%
4.3%
9
Nature Communications
5641 papers in training set
Top 38%
2.7%
10
Communications Biology
993 papers in training set
Top 14%
1.7%
11
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 28%
1.7%
12
Cell Systems
201 papers in training set
Top 3%
1.7%
13
Frontiers in Immunology
638 papers in training set
Top 7%
1.4%
14
eLife
5828 papers in training set
Top 55%
1.3%
15
Nature Methods
385 papers in training set
Top 5%
1.1%
16
Communications Chemistry
48 papers in training set
Top 1%
1.1%
17
PLOS ONE
5266 papers in training set
Top 55%
1.1%
18
ImmunoInformatics
12 papers in training set
Top 0.2%
1.0%
19
Nature Machine Intelligence
70 papers in training set
Top 2%
0.9%
20
Nature Structural & Molecular Biology
18 papers in training set
Top 0.4%
0.9%
21
Structure
193 papers in training set
Top 2%
0.8%
22
Scientific Reports
3612 papers in training set
Top 79%
0.6%
23
Journal of Molecular Biology
232 papers in training set
Top 4%
0.6%