Sample-efficient Antibody Design through Protein Language Model for Risk-aware Batch Bayesian Optimization
Wang, Y.; Wang, B.; Shi, T.; Fu, J.; Zhou, Y.; Zhang, Z.
Show abstract
Antibody design is a time-consuming and expensive process that often requires extensive experimentation to identify the best candidates. To address this challenge, we propose an efficient and risk-aware antibody design framework that leverages protein language models (PLMs) and batch Bayesian optimization (BO). Our framework utilizes the generative power of protein language models to predict candidate sequences with higher naturalness and a Bayesian optimization algorithm to iteratively explore the sequence space and identify the most promising candidates. To further improve the efficiency of the search process, we introduce a risk-aware approach that balances exploration and exploitation by incorporating uncertainty estimates into the acquisition function of the Bayesian optimization algorithm. We demonstrate the effectiveness of our approach through experiments on several benchmark datasets, showing that our framework outperforms state-of-the-art methods in terms of both efficiency and quality of the designed sequences. Our framework has the potential to accelerate the discovery of new antibodies and reduce the cost and time required for antibody design.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GexMolGen: Cross-modal Generation of Hit-like Molecules via Large Language Model Encoding of Gene Expression Signatures 95%
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 95%
- BatchDTA: Implicit batch alignment enhances deep learning-based drug-target affinity estimation 95%
Similar papers in this journal
- GOBoost: Leveraging Long-Tail Gene Ontology Terms for Accurate Protein Function Prediction 94%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 94%
- Attention-based approach to predict drug-target interactions across seven target superfamilies 94%
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 94%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 94%
- Predicting RNA Sequence-Structure Likelihood via Structure-Aware Deep Learning 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.