Back

Improved multimodal protein language model-driven universal biomolecules-binding protein design with EiRA

Zeng, W.; Zou, H.; Li, X.; Wang, X.; Peng, S.

2025-09-05 bioinformatics
10.1101/2025.09.02.673615 bioRxiv
Show abstract

The interactions between proteins and biomolecules form a complex system that supports life activities. Designing proteins capable of targeted biomolecular binding is therefore critical for protein engineering and gene therapy. Here, we propose a new generative model, EiRA, specifically designed for universal biomolecular-binding protein design, which undergoes two-stage post-training, i.e., domain-adaptive masking training and binding site-informed preference optimization, based on a general multimodal protein language model. A systemic evaluation reveals the SOTA performance of EiRA, including structural confidence, diversity, novelty, and designability on 8 test sets across 6 biomolecule types. Meanwhile, EiRA provides a better characterization for biomolecular-binding proteins than generic model, thereby improving the predictive performance of various downstream tasks. We also mitigate severe repetition generation in the original language model by optimizing training strategies and loss. Additionally, we introduced DNA information into EiRA to support DNA-conditioned binder design, further expanding the boundaries of the design paradigm. Purification experiments and molecular dynamics simulations verified the manufacturability and DNA-binding ability of the designed highly differentiated protein. Remarkably, EiRA achieved the "one-shot" design of a Glucagon peptide binder with SPR-confirmed micromolar affinity.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.