Back

A leakage-controlled benchmark shows apparent codon-language-model advantages in synonymous-variant prediction are evaluation artifacts

Liang, Y.; Zhu, W.; Liang, H.; Pan, X.

2026-08-18 bioinformatics
10.64898/2026.08.12.744371 bioRxiv
Show abstract

Synonymous codon choices shape mRNA stability, translation, and folding, and codon language models (cLMs) are increasingly reported to read this biology from sequence. However, when a true signal is thin relative to a confounding one, standard evaluation protocols can manufacture the reported gain rather than measure it--and we show this is what has happened for cLMs on synonymous-variant prediction. Under random splits, the codon advantage is large: tokenization gaps of +2.3-14.3 percentage points (pp) and pretrained codon leads of +2.9 pp over the strongest protein model (ESM-1b) and up to +4.9 pp over ESM-2. We find these numbers are properties of the measurement, not the models. A memorization baseline outscores every neural model; the advantage collapses under gene-held-out evaluation; the sole surviving residual dissolves into six defensible probe defaults; and the synonym-randomization drop that appeared to confirm true signal is itself variance under pooled analysis (0.3 pp, p = 0.49). No advantage survives leakage-controlled evaluation with pooled statistics. We release CodonBench, a leakage-controlled benchmark with an emergent audit cascade, and characterize how artifacts accumulate at every pipeline step. A thin signal (I({sigma}; Y |A) {approx} 0.04 bits) may exist but is not reliably detectable at current sample sizes; we specify what detecting it would require.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.