Efficient Protein Engineering via Integrated Language Models and Bayesian Optimization
Meehl, J.; Siddavatam, P.
Show abstract
This study investigates the application of advanced predictive models to reduce the cost and effort associated with protein engineering campaigns. We explore the use of protein language models (PLMs), a variant of large language models (LLMs), to predict functional performance from protein sequences. A common challenge in this domain is the scarcity of functional data. To address this, we examine zero-shot and few-shot learning methods. Another challenge is efficiently searching the vast fitness landscape for superior protein variants. We evaluate search methods, such as Bayesian optimization, to tackle this problem. The proposed methods are evaluated against a benchmark of 34 protein datasets containing sequences and their quantified functional values. Our findings demonstrate the potential of these advanced predictive models to streamline and accelerate the protein engineering process.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Designing diverse and high-performance proteins with a large language model in the loop 97%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 96%
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 96%
Similar papers in this journal
Similar papers in this journal
- Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-Human interactions 93%
- HiC-GNN: A Generalizable Model for 3D Chromosome Reconstruction Using Graph Convolutional Neural Networks 93%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.