Structure-conditioned masked language models for protein sequence design generalize beyond the native sequence space
Akpinaroglu, D.; Seki, K.; Guo, A.; Zhu, E.; Kelly, M. J.; Kortemme, T.
Show abstract
Machine learning has revolutionized computational protein design, enabling significant progress in protein backbone generation and sequence design. Here, we introduce Frame2seq, a structure-conditioned masked language model for protein sequence design. Frame2seq generates sequences in a single pass, achieves 49.1% sequence recovery on the CATH 4.2 test dataset, and accurately estimates the error in its own predictions, outperforming the autoregressive ProteinMPNN model with over six times faster inference. To probe the ability of Frame2seq to generate novel designs beyond the native-like sequence space it was trained on, we experimentally test 26 Frame2seq designs for de novo backbones with low identity to the starting sequences. We show that Frame2seq successfully designs soluble (22/26), monomeric, folded, and stable proteins (17/26), including a design with 0% sequence identity to native. The speed and accuracy of Frame2seq will accelerate exploration of novel sequence space across diverse design tasks, including challenging applications such as multi-objective optimization.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Improved protein structure refinement guided by deep learning based accuracy estimation 96%
- Hierarchical design of multi-scale protein complexes by combinatorial assembly of oligomeric helical bundle and repeat protein building blocks 96%
- Context-aware geometric deep learning for protein sequence design 96%
Similar papers in this journal
- Efficient protein structure generation with sparse denoising models 96%
- PepFlow: direct conformational sampling from peptide energy landscapes through hypernetwork-conditioned diffusion 96%
- PSICHIC: physicochemical graph neural network for learning protein-ligand interaction fingerprints from sequence data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.