Back

Efficient exploration of sequence space enables rapid generation of functional genome editors

Hughes, N. W.; Kulkarni, S.; Goldman, G.; Marsiglia, J.; Jain, S.; Spees, K.; Hua Fu, B. X.; Vaalavirta, K.; Nakamura, M.

2026-08-20 synthetic biology
10.64898/2026.08.16.745112 bioRxiv
Show abstract

The problem of how protein sequences translate into defined functions remains largely unsolved despite decades of progress. New methods to efficiently explore protein sequence space will help to shed light on these sequence-function relationships, particularly for complex protein function. Here, we describe an approach to create novel, functional proteins through the integration of deep mutational scanning, structural analysis, and evolutionary mining within prompts for a generative protein language model (PLM). We demonstrate the utility of this approach with the generation of novel compact RNA-guided nucleases. This approach is highly efficient, resulting in active nucleases with [~]40% sequence divergence relative to natural proteins and activity equivalent to or exceeding by up to [~]3X that of other compact nucleases at multiple endogenous loci in human cells. The approach described here is rapidly deployable and produces new sequences that will serve as scaffolds for further exploration of complex protein functionality, as well as substrates for novel genome engineering applications.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.