Machine Learning-Driven Optimization of Specific, Compact, and Efficient Base Editors via Single-Round Diversification
Ielanskyi, M.; Wang, M.; Scott, L.; Rieber, L.; Merrett, S.; Schimunek, J.; Mayr, A.; McDowell, I.; Klambauer, G.; Bowen, T.
Show abstract
Cytosine and adenosine base editors show great potential in research and clinical applications. Current iterations of the deaminase--the enzyme used to create precise single-nucleotide changes via base editing--exhibit various off-target effects, including Cas-independent off-targeting, off-base editing, and bystander editing. Engineered deaminases are typically derived from eukaryotic deaminases, which are larger and exhibit high levels of Cas-independent DNA editing, or from evolved variants of the E. coli TadA protein (ecTadA), which are smaller but frequently cause off-base editing. To overcome the limitations inherent to using a single protein sequence as the basis for engineering, we diversified 95 newly identified TadA orthologs by introducing literature-derived mutations and DNA shuffling to yield millions of training sequences for measuring base editor efficiency. Rather than pursuing multiple rounds of random mutagenesis and selection, we trained generative models on the performance data from the diversified pools of variants and drew on information-theoretic insights to efficiently explore the deaminase sequence space to generate diverse and high-performing deaminases. From a single round of diversification, we created a small set of novel and specific cytosine and adenosine deaminases that were markedly distinct in sequence from published base editor deaminases. We additionally found that the deaminases created by our model generally outperform those which we identified through typical directed evolution. The novel adenosine and cytosine deaminases identified in this work showed high on-base activity, comparable to the leading published base editors, but with demonstrably lower off-base activity. The cytosine deaminases were particularly compact compared to known sequences due to a truncation in their final -helix.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A generalizable Cas9/sgRNA prediction model using machine transfer learning with small high-quality datasets 96%
- Guide RNA structure design enables combinatorial CRISPRa programs for biosynthetic profiling 96%
- Genome-wide functional screens enable the prediction of high activity CRISPR-Cas9 and -Cas12a guides in Yarrowia lipolytica 96%
Similar papers in this journal
- The Effect of Pseudoknot Base Pairing on Cotranscriptional Structural Switching of the Fluoride Riboswitch 95%
- Quantum biological insights into CRISPR-Cas9 sgRNA efficiency from explainable-AI driven feature engineering 95%
- Live-cell imaging reveals the trade-off between target search flexibility and efficiency for Cas9 and Cas12a 95%
Similar papers in this journal
Similar papers in this journal
- Yatakemycin biosynthesis requires two deoxyribonucleases for toxin self-resistance 94%
- Photoaffinity enabled transcriptome-wide identification of splice modulating small molecule-RNA binding events in native cells 93%
- Molecular characterisation of the acyltransferase-acyl carrier protein interface in a fungal highly reducing polyketide synthase. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.