Back

Carbon: Decoding the Language of Life

Allal, L. B.; Li, Q.; Fiusco, M.; Tunstall, L.; Rasul, K.; Beeching, E.; Aubakirova, D.; Patino, C.; Frere, T.; Lozhkov, A.; Channing, G.; Wolf, T.; Bernardo, D. d.; Werra, L. v.

2026-05-25 genomics
10.64898/2026.05.22.727119 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWGenomic foundation models have emerged alongside the rapid progress of large language models, offering a promising framework for learning general-purpose sequence priors for DNA understanding, generation, and design. This connection to LLMs creates a major opportunity: modern architectures, scaling infrastructure, autoregressive training, and token-based modeling provide powerful tools for genomic sequence modeling. At the same time, DNA differs fundamentally from natural language. Genomic sequences are noisy, redundant, sparsely constrained, unevenly annotated, and shaped by evolutionary rather than communicative pressures. As a result, key components of the standard LLM recipe, including data construction, tokenization, and training objectives, must be reconsidered in the biological sequence setting. A central challenge in DNA modeling is reconciling single-nucleotide resolution with long-context reasoning. Single-nucleotide resolution is essential for variant effect prediction, splice-site analysis, and codon-level reasoning. Long-context modeling is equally important, as many genomic mechanisms depend on distal regulatory elements, gene neighborhoods, and long-range evolutionary constraints. However, the most direct path to nucleotide-level reasoning, single-nucleotide tokenization, makes genomic sequences extremely long and imposes substantial computational cost on Transformer models. We present CO_SCPLOWARBONC_SCPLOW, a family of efficient generative DNA language models designed as a practical reference point for this setting. CO_SCPLOWARBONC_SCPLOW includes 3B- and 8B-parameter decoder-only autoregressive models using non-overlapping 6-mer tokenization. CO_SCPLOWARBONC_SCPLOW-3B supports a maximum context length of 65,536 tokens, corresponding to approximately 393 kbp of DNA; CO_SCPLOWARBONC_SCPLOW-8B supports up to 131,072 tokens, roughly 786 kbp. This simple and controlled setup helps isolate a central question for DNA language modeling: whether current progress is limited primarily by model architecture and nominal context length, or by more basic alignment between data, tokenization, objectives, evaluation, and the biological structure of genomic sequence. In our training-free evaluation suite, CO_SCPLOWARBONC_SCPLOW-3B is competitive with Evo2-7B despite having less than half the parameters. CO_SCPLOWARBONC_SCPLOW-8B improves on CO_SCPLOWARBONC_SCPLOW-3B on every training-free task, with the largest gain on long-context retrieval. Both models deliver tens-fold faster inference under comparable settings. The CO_SCPLOWARBONC_SCPLOW recipe combines annotation-aware data curation, deterministic 6-mer tokenization, and a staged CE-to-FNS objective schedule, adapting the LLM recipe to the statistical and biological properties of DNA rather than directly transplanting it. We release the models, data, training code, and evaluation suite, including new training-free probes for sequence-level perturbation and DNA long-context retrieval. CO_SCPLOWARBONC_SCPLOW is intended as an open recipe for efficient generative DNA modeling rather than an argument for any specific architecture, tokenization strategy, or objective design as the optimal solution. Its strong performance provides grounded evidence that substantial room remains for domain-aware model design carefully aligned with the genomic sequence itself.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Genome Research
468 papers in training set
Top 0.2%
14.6%
2
Nature Biotechnology
172 papers in training set
Top 0.2%
10.7%
3
Genome Biology
637 papers in training set
Top 1%
9.5%
4
Nucleic Acids Research
1281 papers in training set
Top 2%
7.6%
5
Nature Methods
385 papers in training set
Top 1%
7.6%
50% of probability mass above
6
Bioinformatics
1204 papers in training set
Top 4%
5.3%
7
Nature Communications
5641 papers in training set
Top 28%
5.3%
8
Nature
645 papers in training set
Top 3%
5.3%
9
Nature Genetics
286 papers in training set
Top 2%
3.9%
10
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.3%
11
PLOS Computational Biology
1863 papers in training set
Top 11%
3.0%
12
Nature Computational Science
55 papers in training set
Top 0.5%
2.1%
13
Nature Machine Intelligence
70 papers in training set
Top 2%
1.6%
14
Molecular Biology and Evolution
542 papers in training set
Top 3%
1.6%
15
Science
477 papers in training set
Top 6%
1.5%
16
BMC Bioinformatics
457 papers in training set
Top 5%
1.3%
17
eLife
5828 papers in training set
Top 59%
1.1%
18
Cell
431 papers in training set
Top 8%
1.1%
19
Scientific Reports
3612 papers in training set
Top 69%
1.0%
20
Cell Genomics
172 papers in training set
Top 4%
1.0%
21
GENETICS
483 papers in training set
Top 5%
0.8%
22
Bioinformatics Advances
203 papers in training set
Top 5%
0.8%
23
PLOS ONE
5266 papers in training set
Top 66%
0.6%