Function-guided protein design by deep manifold sampling
Gligorijevic, V.; Berenberg, D.; Ra, S.; Watkins, A.; Kelow, S.; Cho, K.; Bonneau, R.
Show abstract
Protein design is challenging because it requires searching through a vast combinatorial space that is only sparsely functional. Self-supervised learning approaches offer the potential to navigate through this space more effectively and thereby accelerate protein engineering. We introduce a sequence denoising autoencoder (DAE) that learns the manifold of protein sequences from a large amount of potentially unlabelled proteins. This DAE is combined with a function predictor that guides sampling towards sequences with higher levels of desired functions. We train the sequence DAE on more than 20M unlabeled protein sequences spanning many evolutionarily diverse protein families and train the function predictor on approximately 0.5M sequences with known function labels. At test time, we sample from the model by iteratively denoising a sequence while exploiting the gradients from the function predictor. We present a few preliminary case studies of protein design that demonstrate the effectiveness of this proposed approach, which we refer to as "deep manifold sampling", including metal binding site addition, function-preserving diversification, and global fold change.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Computational design of novel Cas9 PAM-interacting domains using evolution-based modelling and structural quality assessment 97%
- Designing diverse and high-performance proteins with a large language model in the loop 96%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 95%
Similar papers in this journal
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 95%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 95%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 94%
Similar papers in this journal
- ECloudGen: Leveraging Electron Clouds as a Latent Variable to Scale Up Structure-based Molecular Design 95%
- Adversarial domain translation networks for fast and accurate integration of large-scale atlas-level single-cell datasets 94%
- Automated customization of large-scale spiking network models to neuronal population activity 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.