Deep Evolutionary Fitness Inference for Variant Nomination from Directed Evolution
Shen, M. W.; Diamant, N.; Helmling, C.; Newland, R.; Lu, Z.; Fannjiang, C.; Kelow, S.; Frey, N.; Saremi, S.; Kelly, R.; Bonneau, R.; Scalia, G.; Cunningham, C.; Biancalani, T.
Show abstract
Iterative screening techniques, such as directed evolution, enable high-throughput affinity maturation to optimize binders to molecular interfaces. However, the decision problem of selecting variants from rich, evolved populations to enter low-throughput follow-up methods remains a significant bottleneck. Here, we present evolutionary fitness inference (EVFI) and DeepEVFI, two machine learning methods that model directed evolution from time-series sequencing data, and infer fitness, a variants ability to enrich under selection pressure. Our methods flexibly handle mutation mechanisms and starting populations that may be partially unknown - settings relevant to drug discovery - and achieve strong performance on a diverse set of experimental data. We conducted two experimental directed evolution campaigns, using antibodies and macrocyclic peptides libraries to identify and optimize binders to therapeutically relevant targets. EVFI and DeepEVFI identified tighter binders that were missed by human experts using conventional frequency-based approaches, including "rising stars" with low frequency. Beyond initial hit discovery, EVFI and Deep-EVFI enables labeling large-scale sequence-fitness datasets and identifying variants of initial binders with diverse properties.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Efficient evolution of human antibodies from general protein language models and sequence information alone 96%
- Model-directed generation of CRISPR-Cas13a guide RNAs designs artificial sequences that improve nucleic acid detection 95%
- Probing molecular specificity with deep sequencing and biophysically interpretable machine learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.