Adapting protein language models for structure-conditioned design
Ruffolo, J. A.; Bhatnagar, A.; Beazer, J.; Nayfach, S.; Russ, J.; Hill, E.; Hussain, R.; Gallagher, J.; Madani, A.
Show abstract
Generative models for protein design trained on experimentally determined structures have proven useful for a variety of design tasks. However, such methods are limited by the quantity and diversity of structures used for training, which represent a small, biased fraction of protein space. Here, we describe proseLM, a method for protein sequence design based on adaptation of protein language models to incorporate structural and functional context. We show that proseLM benefits from the scaling trends of underlying language models, and that the addition of non-protein context - nucleic acids, ligands, and ions - improves recovery of native residues during design by 4-5% across model scales. These improvements are most pronounced for residues that directly interface with non-protein context, which are faithfully recovered at rates >70% by the most capable proseLM models. We experimentally validated proseLM by optimizing the editing efficiency of genome editors in human cells, achieving a 50% increase in base editing activity, and by redesigning therapeutic antibodies, resulting in a PD-1 binder with 2.2 nM affinity.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Synthetic coevolution reveals adaptive mutational trajectories of neutralizing antibodies and SARS-CoV-2 97%
- Single-sequence protein-RNA complex structure prediction by geometric attention-enabled pairing of biological language models 96%
- Undersampling and the inference of coevolution in proteins 95%
Similar papers in this journal
- AlphaFold2 has more to learn about protein energy landscapes 96%
- A multiscale functional map of somatic mutations in cancer integrating protein structure and network topology 95%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.