A sequence-to-function model to predict T7 transcription rates and redesign T7 expression systems with lowered production of immunogenic RNA byproducts
McLellan, J. R.; Salis, H. M.
Show abstract
T7 RNA polymerase is widely used to produce RNA using a canonical T7 promoter; however, it will also bind to low-affinity sites to generate cryptic transcription and produce RNA byproducts, which reduce full-length mRNA purity and yield. When manufacturing therapeutic RNAs for clinical applications, RNA byproducts must be removed using costly downstream purification and can cause adverse immunogenicity. To predict T7 transcription rates and reduce cryptic transcription, we designed 11588 T7 promoters and measured their mRNA levels, spanning a 6300-fold range within in vitro transcription reactions. We developed the T7 Promoter Calculator, a sequence-to-function machine learning model that predicts the T7 transcription rate on arbitrary DNA sequence across a 500-fold range with high accuracy (R2 = 0.80), accounting for both core and flanking motif sequences. We combined the model with generative design to remove low-affinity T7 sites from a therapeutic T7 expression system, resulting in a 2-fold increase in full-length mRNA purity. The automated design of T7 expression systems to remove undesired RNA byproducts increases mRNA purity and lowers downstream separation costs, while reducing adverse immunogenicity.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Deep Generative Optimization of mRNA Codon Sequences for Enhanced mRNA Translation and Therapeutic Efficacy 94%
- High-Throughput 5' UTR Engineering for Enhanced Protein Production in Non-Viral Gene Therapies 94%
- Biochemical-free enrichment or depletion of RNA classes in real-time during direct RNA sequencing with RISER 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.