Back

Maintaining transcriptome solubility constrains mRNA sequence composition

Todisco, M.; Ausler, C.; Jain, A.

2026-05-22 biophysics
10.1101/2025.09.19.677370 bioRxiv
Show abstract

RNA is built from a four-nucleotide alphabet. Complementary sequences inevitably arise, creating pervasive opportunities for promiscuous RNA-RNA interactions. Here, we show that this chemistry makes the transcriptome intrinsically prone to self-association. We simulated the simultaneous interactions of ~7,500 mRNAs representing the E. coli transcriptome at physiological concentrations. These large-scale simulations predict widespread, dynamic clustering driven by RNA alone and organized by long, multivalent transcripts. Purified mRNA recapitulates this behavior in vitro, with aggregate composition mirroring model predictions. Strikingly, native mRNA sequences are markedly less prone to self-association than matched randomized controls: they fold more stably, expose shorter single-stranded regions, and form weaker intermolecular contacts. Similar signatures are observed in abundant human mRNAs, suggesting that evolution has shaped coding sequences to minimize self-association. These findings identify transcriptome solubility as an unrecognized constraint on mRNA sequence evolution and provide a framework for understanding how cells keep their transcriptomes dispersed and functional.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.