Back

Antibody Library Design by Seeding Linear Programming with Inverse Folding and Protein Language Models

Hayes, C. F.; Magana-Zook, S. A.; Goncalves, A.; Solak, A. C.; Faissol, D.; Landajuela, M.

2025-01-23 bioinformatics
10.1101/2024.11.03.621763 bioRxiv
Show abstract

Designing effective antibody libraries is a challenging combinatorial search problem in computational biology. We propose a novel integer linear programming (ILP) method that explicitly controls diversity and affinity objectives when generating candidate libraries. Our approach formulates library design as a constrained optimization problem, where diversity parameters and predicted binding scores are encoded as ILP constraints and objectives. Predicted binding scores are obtained via AI-guided mutational fitness profiling, which combines protein language models and inverse folding tools to evaluate mutational effects. We demonstrate the method on coldstart design tasks for Trastuzumab, D44.1, and Spesolimab, showing that our optimized libraries outperform baseline designs in both predicted affinity and sequence diversity. This hybrid search-and-learning framework illustrates how constrained optimization and predictive modeling can be combined to deliver interpretable, high-quality solutions to antibody library engineering. Code is available at https://github.com/llnl/protlib-designer. ACM Reference FormatConor F. Hayes, Andre R. Goncalves, Steven Magana-Zook, Jacob Pettit, Ahmet Can Solak, Daniel Faissol, and Mikel Landajuela. 2026. Combinatorial Optimization of Antibody Libraries via Constrained Integer Programming. In Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), Paphos, Cyprus, May 25 - 29, 2026, IFAAMAS, 22 pages.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.