TOM: Transformer Optimization of mRNA initiation
Gupta, R.
Show abstract
Recombinant Protein Production enables scientists to insert custom DNA into host cells to produce specific proteins. Current industry standard tools to optimize recombinant DNA use the Codon Adaptation Index (CAI), yet doing this changes local mRNA secondary structure and Minimum Free Energy (MFE) at the translation initiation region, creating hairpins that block ribosome loading and initiation, thus limiting protein production. mRNA secondary structure around the start codon (disrupting ribosome docking/initiation) is the true rate limiting barrier to protein synthesis, and its 22x more correlated to protein production than CAI. This project introduces TOM, a novel Transformer deep learning model that optimizes mRNA secondary structure at the critical translation initiation region. Over 40 million natural E. Coli initiation sequences were downloaded, where through a rigorous data filtration effort, only 10,000 ideal, non redundant, and naturally occurring sequences were used to train the model. TOM is benchmarked on MFE, adenine count, codon usage, and a negative element analysis against industry standard optimizers and significantly outperforms. TOMs improvement of the MFE at the initiation region offers a significant increase in protein production by addressing the rate-limiting step of translation initiation.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- ICOR: Improving codon optimization with recurrent neural networks 96%
- CysPresso: A classification model utilizing deep learning protein representations to predict recombinant expression of cysteine-dense peptides 95%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 94%
Similar papers in this journal
Similar papers in this journal
- LMPred: Predicting Antimicrobial Peptides Using Pre-Trained Language Models and Deep Learning 95%
- UTRGAN: Learning to Generate 5' UTR Sequences for Optimized Translation Efficiency and Gene Expression 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.