Reading TEA leaves for de novo protein design
Pantolini, L.; Durairaj, J.
Show abstract
De novo protein design expands the functional protein universe beyond natural evolution, offering vast therapeutic and industrial potential. Monte Carlo sampling in protein design is under-explored due to the typically long simulation times required or prohibitive time requirements of current structure prediction oracles. Here we make use of a 20-letter structure-inspired alphabet derived from protein language model embeddings to score random mutagenesis-based Metropolis sampling of amino acid sequences. This facilitates fast template-guided and unconditional design, generating sequences that satisfy in silico designability criteria without known homologues. Ultimately, this unlocks a new path to fast and de novo protein design.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ProtMamba: a homology-aware but alignment-free protein state space model 96%
- Multi-Scale Structural Analysis of Proteins by Deep Semantic Segmentation 96%
- QDeep: distance-based protein model quality estimation by residue-level ensemble error classifications using stacked deep residual neural networks 96%
Similar papers in this journal
- De novo protein design by inversion of the AlphaFold structure prediction network 97%
- Neural Network-Derived Potts Models for Structure-Based Protein Design using Backbone Atomic Coordinates and Tertiary Motifs 97%
- Enriching stabilizing mutations through automated analysis of molecular dynamics simulations using BoostMut 96%
Similar papers in this journal
- To pack or not to pack: revisiting protein side-chain packing in the post-AlphaFold era 95%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 94%
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.