Back

Antibody-Antigen Affinity Prediction with Chain-Aware Protein Language Modeling

Singh, H.; Malhotra, A.; Srivastava, S. P.; SINGH, R. K.; Gorantla, R.

2026-06-21 bioinformatics
10.64898/2026.06.19.733375 bioRxiv
Show abstract

MotivationAntibody-antigen affinity determines which antibodies advance in therapeutic discovery, repertoire analysis and affinity maturation, but experimental measurements are sparse relative to the scale of sequence libraries. Structure-based predictors can exploit interface geometry when reliable complexes are available, yet early discovery often requires ranking many heavy-light chain pairs against antigens for which no complex structure exists. Existing sequence-based models are scalable, but frequently compress heavy and light chains into a single antibody representation or concatenate antibody and antigen features obscuring the chain-specific and epitope-specific signals that drive binding. ResultsWe present AbAffinity, a sequence-only chain-aware three-stream architecture that maintains heavy chain, light chain and antigen as distinct streams. It integrates frozen ESM-2 embeddings with heavy-chain CDR-focused pooling, heavy-light self-attention, adaptive fusion gating and gated cross-attention, training only a compact interaction module. On the SAAINT-DB benchmark, AbAffinity achieves strong predictive performance under ten-fold cross-validation and maintains robust accuracy on novel antigens. It consistently outperforms recent sequence-based models across external benchmarks including SAbDab, AB-Bind and SKEMPI 2.0. Ablation studies highlight the contributions of chain-specific representations, CDR-focused pooling and the gated interaction pathway. Integrated Gradients attributions recover known paratope and epitope residues at structurally validated interfaces. AbAffinity provides a lightweight, explainable sequence-first framework for antibody triage and prioritisation when structural information is limited or unavailable.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
15.0%
2
Bioinformatics
1204 papers in training set
Top 2%
12.6%
3
mAbs
32 papers in training set
Top 0.1%
12.6%
4
Nature Communications
5641 papers in training set
Top 18%
9.7%
50% of probability mass above
5
Cell Systems
201 papers in training set
Top 0.8%
5.5%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.8%
7
PLOS Computational Biology
1863 papers in training set
Top 9%
3.5%
8
Nature Methods
385 papers in training set
Top 3%
3.5%
9
Patterns
78 papers in training set
Top 0.7%
2.8%
10
Bioinformatics Advances
203 papers in training set
Top 2%
2.6%
11
Cell Reports Methods
165 papers in training set
Top 1%
2.0%
12
Journal of Chemical Information and Modeling
238 papers in training set
Top 2%
1.7%
13
Molecular Systems Biology
162 papers in training set
Top 2%
1.5%
14
Nature Biotechnology
172 papers in training set
Top 3%
1.4%
15
Nature
645 papers in training set
Top 8%
1.1%
16
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
17
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
18
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 37%
1.1%
19
iScience
1154 papers in training set
Top 30%
1.0%
20
Nature Computational Science
55 papers in training set
Top 1%
1.0%
21
BMC Bioinformatics
457 papers in training set
Top 5%
1.0%
22
Protein Science
246 papers in training set
Top 4%
0.8%
23
Scientific Reports
3612 papers in training set
Top 74%
0.8%
24
npj Systems Biology and Applications
125 papers in training set
Top 2%
0.6%
25
PRX Life
42 papers in training set
Top 1%
0.6%
26
Advanced Science
286 papers in training set
Top 11%
0.6%
27
Genome Biology
637 papers in training set
Top 9%
0.6%