Back

Mapping the Inter- and Intra-genic Codon Usage Landscape in Homo sapiens

Arshad, M.; Uchmanowicz, M.; Rana, V.; Rafiq, M. A.

2025-07-31 genomics
10.1101/2025.07.27.667039 bioRxiv
Show abstract

Although the genetic code is degenerate, codon selection is non-random and reflects significant functional constraints. Codon usage bias (CUB) acts as a layer of post-transcriptional regulation, influencing mRNA stability, translation kinetics, and co-translational protein folding. While CUB is well-characterized in unicellular organisms, its regulatory scope and functional consequences in humans remain complex and less defined. Our study offers a comprehensive evaluation of human codon usage. We report that genes exhibiting the strongest codon bias are enriched in high-stoichiometry biological processes, such as skin development and oxygen/carbon dioxide transport, and harbor significantly fewer synonymous variants than expected ({rho} = -0.24, p < 2.2 x 10-16). Furthermore, we find that codon optimization is spatially regulated: it is significantly more pronounced in structured protein domains compared to intrinsically disordered regions (IDRs) (Cliffs {Delta}= 0.26, p < 2.2 x 10-16). Consistent with translational selection, the most frequently used codons are supported by higher tRNA gene copy numbers ({rho} = 0.49, p < 6.4 x 10-4). Finally, by correcting for GC3 content, we reveal that the apparent correlation between ENC and adaptation indices (CAI/tAI) vanishes, allowing us to disentangle mutational pressure from translational selection. Collectively, our findings position codon usage bias as a central, evolutionarily conserved regulator of translation and protein folding in humans. Our results provide a comprehensive and integrated view of intergenic and intragenic codon usage bias in humans, reinforcing the biological relevance of synonymous codon choice in shaping translational dynamics and protein biogenesis. This provides a refined framework for interpreting synonymous variation and guiding functional genomics.

Published in NAR Genomics and Bioinformatics (predicted rank #4) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.