Back

BLOSUM Is All You Learn - Generative Antibody Models Reflect Evolutionary Priors

Ucar, T.; Sormanni, P.

2025-10-27 bioengineering
10.1101/2025.10.26.684652 bioRxiv
Show abstract

Generative models have emerged as powerful tools for antibody sequence design, with recent studies demonstrating that log-likelihood scores from these models can correlate with binding affinity and potentially serve as effective ranking metrics. This raises a fundamental question: why should log-likelihood scores from generative models correlate with binding affinity? In this work, we investigate the biochemical basis of these model-derived log-likelihoods by comparing them with classical evolutionary similarity metrics. We find that BLOSUM similarity scores between designed and parental antibody sequences correlate strongly with measured binding affinity--on par with the predictive performance of a state-of-the-art diffusion-based generative model. Moreover, these BLOSUM scores also align closely with log-likelihoods from multiple generative models, suggesting that such models may be implicitly learning evolutionary priors encoded in substitution matrices. When computed with respect to a known binder, both BLO-SUM scores and log-likelihoods act as approximate measures of sequence distance from that reference. As this distance increases, the likelihood of a candidate being a binder decreases, explaining the observed correlation between these scores and binding affinity. In contrast, similarity scores based on position weight matrices (PWMs) and position-specific scoring matrices (PSSMs), which do not rely on knowledge of the parental sequence, show weaker and less consistent alignment with binding affinity, with performance depending on the background sequence data. Additionally, using consensus sequences in place of parental sequences to compute BLOSUM scores largely eliminates the observed correlation with affinity, underscoring the context-specific nature of the correlations. These findings highlight the potential of interpretable, evolution-inspired metrics to complement generative modeling in anti-body design, offering insights into both model behavior and biological relevance.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 1%
17.0%
2
PLOS Computational Biology
1863 papers in training set
Top 2%
15.1%
3
mAbs
32 papers in training set
Top 0.1%
9.8%
4
Protein Engineering, Design and Selection
15 papers in training set
Top 0.1%
7.9%
5
Cell Systems
201 papers in training set
Top 0.5%
6.7%
50% of probability mass above
6
Computational and Structural Biotechnology Journal
242 papers in training set
Top 0.5%
5.5%
7
Protein Science
246 papers in training set
Top 1%
3.2%
8
Bioinformatics Advances
203 papers in training set
Top 2%
3.2%
9
Nature Communications
5641 papers in training set
Top 39%
2.4%
10
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 22%
2.4%
11
Journal of Open Source Software
25 papers in training set
Top 0.1%
2.4%
12
Briefings in Bioinformatics
354 papers in training set
Top 4%
1.7%
13
Scientific Reports
3612 papers in training set
Top 54%
1.7%
14
eLife
5828 papers in training set
Top 49%
1.7%
15
Cell Reports Methods
165 papers in training set
Top 2%
1.3%
16
Frontiers in Immunology
638 papers in training set
Top 7%
1.1%
17
Journal of Cheminformatics
29 papers in training set
Top 0.5%
1.1%
18
PLOS ONE
5266 papers in training set
Top 55%
1.1%
19
iScience
1154 papers in training set
Top 29%
1.1%
20
ImmunoInformatics
12 papers in training set
Top 0.2%
0.9%
21
Nucleic Acids Research
1281 papers in training set
Top 13%
0.8%
22
Nature Methods
385 papers in training set
Top 6%
0.8%