Back

Why This Code? A Constrained Mapping Framework for the Evolutionary Stability of the Canonical Genetic Code

Lin, R.; Wang, C.; Hu, Y.; Wang, C.

2026-08-28 evolutionary biology
10.64898/2026.08.28.747708 bioRxiv
Show abstract

The canonical genetic code is used by most known forms of life, yet explaining its historical origin and present day functional performance requires comparison with the enormous space of possible codon-to-output assignments. Here, the code is formulated as a hierarchy of constrained mapping problems spanning codeword length, degeneracy composition, synonymous-block partitioning, semantic assignment and a coarse decoder layer. Exact structural analyses identify triplets as a Pareto choice under a fixed-length full-codebook model and show that anonymous degeneracy statistics alone do not explain the canonical profile. Within a fixed canonical block architecture and under specified objective functions, recurrent AAindex-based rule learning contracts the 20! amino-acid assignment space to 2.72 ** 1011 admissible mappings, from which 108 complete codes are sampled. In this screened conditional candidate library, the standard genetic code ranks in the best 0.9749% under the equal-weight three-objective score and in the best 1.801% when accessible replacement diversity is added. Sensitivity analyses show that this position is broad across many, but not all, tested objective weights and aggregation rules. These results describe a conditional multi-objective compromise; they do not establish global optimality, historical inevitability or cellular feasibility of decoder redesign.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.