Generating novel protein sequences using Gibbs sampling of masked language models
Johnson, S. R.; Massie, K.; Monaco, S.; Sayed, Z.
Show abstract
Recently developed language models (LMs) based on deep neural networks have demonstrated the ability to generate fluent natural language text. LMs pre-trained on protein sequences have shown state of the art performance on a variety of downstream tasks. Protein LMs have also been used to generate novel protein sequences. In the present work we use Gibbs sampling of BERT-style LMs, pre-trained on protein sequences using the masked language modeling task, to generate novel protein sequences. We evaluate the quality of the generated sequences by comparing them to natural sequences from the same family. In particular, we focus on proteins from the chorismate mutase type II family, which has been used in previous work as an example target for protein generative models. We find that the Gibbs sampling process on BERT-style models pretrained on millions to billions of protein sequences is able to generate novel sequences that retain key features of related natural sequences. Further, we find that smaller models fine-tuned or trained from scratch on family-specific data are able to equal or surpass the generation quality of large pre-trained models by some metrics. The ability to generate novel natural-like protein sequences could contribute to the development of improved protein therapeutics and protein-catalysts for industrial chemical production.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Designing diverse and high-performance proteins with a large language model in the loop 97%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 97%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 97%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 96%
- Prop3D: A Flexible, Python-based Platform for Machine Learning with Protein Structural Properties and Biophysical Data 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
Similar papers in this journal
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 97%
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 96%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.