Back

DNA Compression with Genomic Language Models: Tokenization, Benchmarking, and an Information-Content Map

Macala, V.; Simecek, P.

2026-06-12 bioinformatics
10.64898/2026.06.10.731316 bioRxiv
Show abstract

Lossless compression and probabilistic sequence modeling are two faces of the same coin: a model that assigns high probability to a sequence can encode it in few bits via arithmetic coding. We exploit this duality to evaluate genomic language models as compressors of DNA, using compression primarily as an objective probe of generative sequence modeling rather than as a deployable storage system. We release DNAGPT2, a family of ten GPT-2-small models pretrained for one epoch on a single A40 using the DNABERT2 multi-species corpus that differ only in byte-pair encoding vocabulary size. Coupled with arithmetic coding, the best model reaches 1.47 bits per base (bpb) on the T2T human genome, fourth in the Cobilab compression benchmark and ahead of every general-purpose compressor. Our results suggest that NLP-style tokenization choices may be suboptimal for DNA: a 32-token BPE vocabulary compresses better than larger vocabularies. We also find that, in this benchmark, published long-context genomic LMs underperform a much shorter-context BPE GPT-2; we discuss in Section 5 that this is not a controlled context-length ablation, since the compared models also differ in architecture, training data, parameter count, and tokenization. Finally, we compute a per-nucleotide information-content map of the human genome and show that exons, introns, intergenic regions, and Alu repeats have statistically distinct information profiles.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 2%
12.8%
2
Genome Biology
637 papers in training set
Top 0.7%
12.0%
3
Genome Research
468 papers in training set
Top 0.5%
7.9%
4
Nature Biotechnology
172 papers in training set
Top 0.4%
6.8%
5
Bioinformatics Advances
203 papers in training set
Top 1%
4.1%
6
Nature Communications
5641 papers in training set
Top 32%
4.1%
7
Nature Methods
385 papers in training set
Top 2%
4.1%
50% of probability mass above
8
NAR Genomics and Bioinformatics
242 papers in training set
Top 1%
3.2%
9
Nucleic Acids Research
1281 papers in training set
Top 7%
2.7%
10
Algorithms for Molecular Biology
17 papers in training set
Top 0.1%
2.5%
11
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.2%
2.5%
12
Scientific Reports
3612 papers in training set
Top 47%
2.1%
13
BMC Bioinformatics
457 papers in training set
Top 3%
2.1%
14
Cell Systems
201 papers in training set
Top 2%
2.0%
15
Journal of Computational Biology
48 papers in training set
Top 0.6%
1.7%
16
PLOS Computational Biology
1863 papers in training set
Top 14%
1.7%
17
Nature Machine Intelligence
70 papers in training set
Top 1%
1.7%
18
Briefings in Bioinformatics
354 papers in training set
Top 4%
1.7%
19
BioData Mining
22 papers in training set
Top 0.3%
1.5%
20
GigaScience
212 papers in training set
Top 3%
1.4%
21
Nature Computational Science
55 papers in training set
Top 1.0%
1.1%
22
iScience
1154 papers in training set
Top 25%
1.1%
23
Frontiers in Genetics
230 papers in training set
Top 4%
1.1%
24
Nature Genetics
286 papers in training set
Top 4%
1.1%
25
Nature
645 papers in training set
Top 8%
1.1%
26
Patterns
78 papers in training set
Top 2%
1.0%
27
GENETICS
483 papers in training set
Top 4%
0.9%
28
PLOS ONE
5266 papers in training set
Top 61%
0.9%
29
Computational and Structural Biotechnology Journal
242 papers in training set
Top 7%
0.9%
30
IEEE/ACM Transactions on Computational Biology and Bioinformatics
38 papers in training set
Top 1%
0.6%