Back

CodonBERT and ESM-2 Embedding Spaces Share an Evolutionarily Conserved Paired Geometry Encoding Synonymous Codon Information

Lu, B.

2026-07-03 bioinformatics
10.64898/2026.06.30.735115 bioRxiv
Show abstract

Synonymous codons encode the same amino acid yet are used non-randomly across genomes, a phenomenon with well-documented functional consequences for translation efficiency and mRNA stability. Whether the information embedded in synonymous codon choice is recoverable from the internal representations of independently trained deep learning models remains an open question. Here we show that CodonBERT, a nucleotide language model trained exclusively on coding sequences, and ESM-2, a protein language model, harbor a linearly retrievable paired geometry that persists after removing amino acid composition effects and encodes synonymous-codon-level information. This paired geometry transfers across species, sharpens under ortholog restriction, and decays monotonically with evolutionary distance from rat (AUC 0.96) through mouse and zebrafish (0.92) to yeast (0.73). The signal is specific to codon-aware nucleotide models: DNABERT-2 and Nucleotide Transformer v2 achieve only 0.16 and 0.13 R@1, respectively. These results demonstrate that deep learning models independently trained on distinct molecular modalities converge on evolutionarily constrained representations that capture synonymous codon information.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.