Unexplored regions of the protein sequence-structure map revealed at scale by a library of foldtuned language models
Subramanian, A. M.; Thomson, M.
Show abstract
Amino-acid sequence space is combinatorially vast, with well-folded proteins distributed sparsely and connected by vanishingly few permissible mutational paths. Novel-in-sequence versions of structures observed in nature promise to sample features such as new binding motifs and active site geometries but are rendered inaccessible to evolution or direct search by the extent of sequence perturbations required. Here we introduce a novel algorithm - termed "foldtuning" - that leverages principles of adversarial learning to drive protein language models (PLMs) to erase detectable homology to natural sequences while preserving a target structure, systematically traversing protein-space without being limited by evolutionary barriers. We build foldtuned PLMs for >700 targets including membrane-bound receptors, redox enzymes, and signaling domains. Foldtuned proteins are diverse and far-from-natural in sequence, filling out structurally-equivalent families defined by fundamental biophysical constraints invisible to traditional sequence-based bioinformatics methods. Experimental characterization demonstrates that foldtuned proteins express stably in vitro and function in vivo. By revealing sequence-structure information at scale beyond evolution, foldtuning promises to accelerate the reconstitution and realization of novel-to-nature systems for synthetic biology problems from therapeutics to catalysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Generalizable and scalable protein stability prediction with rewired protein generative models 96%
- Multi-scale classification decodes the complexity of the human E3 ligome 96%
- Hierarchical design of multi-scale protein complexes by combinatorial assembly of oligomeric helical bundle and repeat protein building blocks 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.