Splice-Aware Optimization Prevents Pervasive Missplicing of Natural and Synthetic cDNAs
Ahmad, M.; Bellido Molias, F.; Cano Aroca, L.; Mordstein, C.; Watson, S.; Bhuiyan, N.; Clarke, M. J.; Kimchi-Sarfaty, C.; Katneni, U.; Gaunt, E. R.; Hurst, L. D.; Netuschil, N.; Hofmeister, T.; Liss, M.; Kudla, G.
Show abstract
Heterologous gene expression is widely used across biology and medicine, and often relies on codon optimization to increase protein yields. Here we uncover missplicing as a common and largely unrecognized failure mode of heterologous expression. Using systematically designed libraries comprising over 5,000 synthetic reporter genes and natural human cDNAs, we find that the majority of gene variants expressed in a human cell line are at least partially spliced, and in many variants the spliced isoform dominates, reducing protein output or ablating expression entirely. By analysing sequence determinants of expression across multiple human cell lines, we uncover a hierarchical architecture of regulatory control, where GC content establishes baseline mRNA levels, local sequence features influence splicing, and tissue-specific codon adaptation to tRNA pools fine-tunes translation efficiency. These findings enable us to develop predictive models of expression and splicing, benchmark current optimization strategies, and design a splice-aware optimization algorithm that substantially improves transgene performance.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning-optimized targeted detection of alternative splicing 96%
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 96%
- Ultradeep characterisation of translational sequence determinants refutes rare-codon hypothesis and unveils quadruplet base pairing of initiator tRNA and transcript 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.