Pre-trained protein language model for codon optimization
Pathak, S.; Lin, G.
Show abstract
Messenger ribonucleic acid (mRNA) vaccines represent a major advancement in synthetic biology, yet their efficacy remains limited by how efficiently the encoded protein is translated within the host. Since multiple codons can code for the same amino acid, the search space of possible coding sequences (CDS) within mRNA grows exponentially with protein length, making the problem highly underdetermined. Finding CDSs that yield efficient translation hinges on codon optimization--the process of choosing among synonymous codons, while encoding the same protein but differ in their effects on translation speed, tRNA availability, and mRNA secondary structure. Recent deep learning approaches have framed codon optimization as a sequence learning problem, where the goal is to model context-dependent codon usage patterns across the amino acid in protein sequence. However, these methods rely on large sequence models that learn amino-acid embeddings from scratch, leading to computationally intensive training. We propose ppLM-CO, a lightweight codon optimization framework that integrates pretrained protein language models (ppLMs) to directly provide contextual amino-acid embeddings, thereby eliminating the need for embedding learning. This design reduces trainable parameters by over 92% - 99% compared with prior deep models while maintaining complete biological fidelity. In-silico evaluations across three species and two vaccine targets--SARS-CoV-2 spike and Varicella-Zoster Virus (VZV) gE viral proteins--demonstrate that ppLM-CO consistently achieves higher expression and competitive stability, establishing a scalable and biologically consistent approach for codon optimization.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Predicting RNA Sequence-Structure Likelihood via Structure-Aware Deep Learning 96%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 94%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 96%
- Representation learning applications in biological sequence analysis 95%
- Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-Human interactions 94%
Similar papers in this journal
- Prediction of virus-host association using protein language models and multiple instance learning 95%
- DeepHE: Accurately Predicting Human Essential Genes based on Deep Learning 93%
- MoCETSE: A mixture-of-convolutional experts and transformer-based model for predicting Gram-negative bacterial secreted effectors 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.