Sequence-free landscape inference for directed evolution
Towers, S.; James, J.; Steel, H.; Kempf, I.
Show abstract
Directed evolution is a method for engineering biological systems or components, such as proteins, wherein desired traits are optimised through iterative rounds of mutagenesis and selection of fit variants. The process of protein directed evolution can be envisaged as navigation over high-dimensional landscapes with numerous local maxima. The performance of any strategy in navigating such a landscape is dependent on the ruggedness of that landscape. However, this information is generally unavailable at the outset of an experiment, and cannot currently be computed using analytical methods. Here we propose SLIDE, Sequence-free Landscape Inference for Directed Evolution, which consists of two parts. First, SLIDE provides an estimation for landscape ruggedness from a mutating population using only population-level phenotypic data and an estimation of mutation rate. Ruggedness information in itself is valuable in protein design, for instance in predicting evolutionary stability. Second, SLIDE offers a framework for using the estimated ruggedness metric to select high-performing parameters for directed evolution control. Using theoretical NK landscapes and four real-world protein fitness landscapes, we demonstrate improvement upon the performance of standard selection strategies, particularly on rugged landscapes, using a pipeline that could also be combined with emerging AI-based methods for driving direction evolution.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Inference of annealed protein fitness landscapes with AnnealDCA 96%
- Conflicting effects of recombination on the evolvability and robustness in neutrally evolving populations 96%
- An extension of the Walsh-Hadamard transform to calculate and model epistasis in genetic landscapes of arbitrary shape and complexity 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.