A leakage-controlled benchmark shows apparent codon-language-model advantages in synonymous-variant prediction are evaluation artifacts
Liang, Y.; Zhu, W.; Liang, H.; Pan, X.
Show abstract
Synonymous codon choices shape mRNA stability, translation, and folding, and codon language models (cLMs) are increasingly reported to read this biology from sequence. However, when a true signal is thin relative to a confounding one, standard evaluation protocols can manufacture the reported gain rather than measure it--and we show this is what has happened for cLMs on synonymous-variant prediction. Under random splits, the codon advantage is large: tokenization gaps of +2.3-14.3 percentage points (pp) and pretrained codon leads of +2.9 pp over the strongest protein model (ESM-1b) and up to +4.9 pp over ESM-2. We find these numbers are properties of the measurement, not the models. A memorization baseline outscores every neural model; the advantage collapses under gene-held-out evaluation; the sole surviving residual dissolves into six defensible probe defaults; and the synonym-randomization drop that appeared to confirm true signal is itself variance under pooled analysis (0.3 pp, p = 0.49). No advantage survives leakage-controlled evaluation with pooled statistics. We release CodonBench, a leakage-controlled benchmark with an emergent audit cascade, and characterize how artifacts accumulate at every pipeline step. A thin signal (I({sigma}; Y |A) {approx} 0.04 bits) may exist but is not reliably detectable at current sample sizes; we specify what detecting it would require.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EternaBrain: Automated RNA design through move sets from an Internet-scale RNA videogame 94%
- Improving deep models of protein-coding potential with a Fourier-transform architecture and machine translation task 94%
- PandoGen: Generating complete instances of future SARS-CoV-2 sequences using Deep Learning 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.