DyAb: sequence-based antibody design and property prediction in a low-data regime
Lin, J. Y.-Y.; Hofmann, J. L.; Leaver-Fay, A.; Liang, W.-C.; Vasilaki, S.; Lee, E.; Pinheiro, P. O.; Tagasovska, N.; Kiefer, J. R.; Wu, Y.; Seeger, F.; Bonneau, R.; Gligorijevic, V.; Watkins, A.; Cho, K.; Frey, N. C.
10.1101/2025.01.28.635353 bioRxivShow abstract
Protein therapeutic design and property prediction are frequently hampered by data scarcity. Here we propose a new model, DyAb, that addresses these issues by leveraging a pair-wise representation to predict differences in protein properties, rather than absolute values. DyAb is built on top of a pre-trained protein language model and achieves a Spearman rank correlation of up to 0.85 on binding affinity prediction across molecules targeting three different antigens (EGFR, IL-6, and an internal target), given as few as 100 training data. We employ DyAb in two design contexts: as a ranking model to score combinations of known mutations, and combined with a genetic algorithm to generate new sequences. Our method consistently generates novel antibody candidates with high binding rates, including designs that improve on the binding affinity of the lead molecule by more than ten-fold. DyAb represents a powerful tool for engineering therapeutic protein properties in low data regimes common in early-stage drug development.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sliding Window INteraction Grammar (SWING): a generalized interaction language model for peptide and protein interactions 95%
- Direct prediction of intrinsically disordered protein conformational properties from sequence 95%
- Megabodies expand the nanobody toolkit for protein structure determination by single-particle cryo-EM 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.