CENO: A Genome-Scale World Model for Evolutionary Sequence Interpretation and Programmable Regulatory Design
Ma, M.; Wu, Y.; Chen, X.; Jiang, F.; Lin, P.; Ye, D.; Sun, Y.; Zhang, Y.; Shi, T.; Zhao, Y.; Ouyang, W.; Zhou, B.; Bai, L.; Ren, Y.
Show abstract
DNA encodes biological function across a continuum of sequence scales, from single-nucleotide and motif-level grammar to regulatory neighborhoods, chromatin-scale organization and evolutionary constraint. A useful model of genomes should therefore do more than classify short sequence windows: it should maintain nucleotide-resolution state over long contexts, score counterfactual mutations, condition on homologous sequence evidence and generate candidates that can be evaluated against structural or functional objectives. We define such a system operationally as a genomic world model: a general-purpose generative model of genome sequence space that unifies sequence understanding and sequence design through a shared state and likelihood interface. Here we introduce CENO, a family of long-context generative genomic world models designed to preserve local DNA grammar while extending usable context to regulatory and chromatin scales. CENO combines Mamba sequence-mixing layers, sparse attention layers and mixture-of-experts capacity in a single autoregressive backbone, and is trained at 300M, 600M and 1B parameter scales with a staged curriculum that progresses from 8k-token cross-domain genomic pretraining to 131k- and 1M-token whole-genome long-context continuation. We evaluate CENO under a unified world-model benchmark paradigm spanning retrieval, representation, counterfactual perturbation, reconstruction, evolutionary conditioning and design. CENO retains practical long-context inference and retrieves distal sequence in synthetic assays. In zero-shot long-context analyses, without task-specific fine-tuning, long-context continuation yields annotation- and chromatin-boundary-associated attention patterns and frozen-state representations that generalize across human cell types and mouse cell or tissue settings. To incorporate evolutionary information, we further post-train CENO on packed real multiple-sequence-alignment contexts and score variants by reference-mutant likelihood deltas, improving matched variant-effect prediction and producing evolutionary enrichment signals across species. Complementing these perturbation-based variant tests, we evaluate zero-shot long-sequence generation by partial-gene continuation, asking whether the model can recover withheld gene-scale sequence structure across eukaryotic, bacterial and archaeal species; recovery improves with model scale and later whole-genome long-context training. Finally, we use CENO as the backbone for a cell-type-specific enhancer design workflow in mouse cortex, coupling a CENO-based accessibility oracle with conditional supervised fine-tuning and oracle-guided reinforcement learning. Together, CENO provides a genome-scale sequence world-model framework for sequence interpretation, evolutionary reasoning, gene-scale reconstruction and programmable regulatory sequence generation.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AlphaGenome: advancing regulatory variant effect prediction with a unified DNA sequence model 98%
- Large-scale clinical interpretation of genetic variants using evolutionary data and deep learning 96%
- Massively parallel characterization of transcriptional regulatory elements in three diverse human cell types 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.