Back

Folding scFv--Antigen Complexes at Scale

Shah, R. N.; Ouyang-Zhang, J.; Cohen, Z.; Briglia, M. R.; Zhang, C.; Klivans, A.; Diaz, D. J.

2026-07-03 bioinformatics
10.64898/2026.07.01.730981 bioRxiv
Show abstract

Accurate modeling of antibody-antigen (Ab-Ag) complexes is central to biologic development, yet the reliability and failures of modern Ab-Ag folding pipelines remain poorly characterized. Single-chain variable fragments (scFvs) are therapeutically important antibodies, but large-scale evaluations of structure prediction models on scFv-Ag complexes are largely lacking. We introduce a scalable benchmarking pipeline that generates large ensembles of scFv-Ag structure predictions by cofolding a curated subset of 3,800 Ab-Ag complexes from SAbDab using multiple state-of-the-art models under diverse inference-time settings. The resulting dataset, SCALE (scFv-Ag CompLex Ensembles) includes standardized scFv-Ag sequences and around 200,000 predicted complexes spanning different models, sampling strategies, and auxiliary inputs. Using SCALE, we evaluate model performance in recovering correct scFv-Ag interfaces and assess the ability of existing confidence metrics to select the best structure from prediction ensembles. We find that while confidence scores effectively distinguish easy from hard scFv-Ag complexes, they often fail to identify the highest-quality interface for a given target. Further analysis shows that near-correct interfaces typically appear in ensembles but at low frequency, and inference-time choices like sampling, recycling, and using evolutionary or structural information are crucial for accurate scFv-Ag complex predictions. Dataset and analysis code are available at https://huggingface.co/datasets/ravishah1/SCALE

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
14.7%
2
Cell Systems
201 papers in training set
Top 0.2%
11.6%
3
Nature Communications
5641 papers in training set
Top 19%
9.4%
4
mAbs
32 papers in training set
Top 0.1%
7.7%
5
Nature Methods
385 papers in training set
Top 2%
6.5%
6
Nature Machine Intelligence
70 papers in training set
Top 0.5%
5.3%
50% of probability mass above
7
Nature Biotechnology
172 papers in training set
Top 0.7%
5.3%
8
PLOS Computational Biology
1863 papers in training set
Top 9%
3.4%
9
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
3.1%
10
Bioinformatics Advances
203 papers in training set
Top 2%
2.6%
11
Briefings in Bioinformatics
354 papers in training set
Top 4%
2.1%
12
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 29%
1.7%
13
Cell Reports Methods
165 papers in training set
Top 2%
1.7%
14
Patterns
78 papers in training set
Top 1%
1.6%
15
Nature Genetics
286 papers in training set
Top 3%
1.5%
16
eLife
5828 papers in training set
Top 54%
1.4%
17
Protein Science
246 papers in training set
Top 3%
1.4%
18
Science
477 papers in training set
Top 7%
1.1%
19
Communications Biology
993 papers in training set
Top 23%
1.1%
20
Nature Computational Science
55 papers in training set
Top 1%
1.1%
21
Nucleic Acids Research
1281 papers in training set
Top 12%
1.0%
22
Frontiers in Immunology
638 papers in training set
Top 9%
1.0%
23
Scientific Reports
3612 papers in training set
Top 72%
0.9%
24
Nature
645 papers in training set
Top 10%
0.9%
25
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.8%
26
Cell Reports
1498 papers in training set
Top 28%
0.8%
27
Molecular Systems Biology
162 papers in training set
Top 4%
0.6%
28
Structure
193 papers in training set
Top 3%
0.6%
29
iScience
1154 papers in training set
Top 42%
0.6%
30
Neuron
337 papers in training set
Top 6%
0.6%