Back

Metagenomic contextualization of proteins with state space models

Azbijari, N.; Wynne, J. H.; David, M.; Thurber, A. R.

2026-07-11 bioinformatics
10.64898/2026.07.07.736993 bioRxiv
Show abstract

Since the early adoption of metagenomics (the culture-free sequencing of microbial community genomes) in 2011, sequence data has increased over 500-fold across ecosystems. This surge in data has outpaced reliable taxonomic and functional annotation, with over half of sequences lacking confident functional assignment. These unknown sequences limit our understanding of microbial processes central to planetary health and human health. Recent advances in genomic language modeling have made progress in the interpretation of metagenomics datasets. Most state-of-the-art models rely on transformer architectures, which limit the maximum sequence length and therefore capture only a fraction of assembled metagenomic sequences due to the quadratic scaling of attention. This prevents training and inference on sequences with broad context, including multiple coding and non-coding regions. To overcome this limitation, we propose leveraging new model architectures that scale linearly with sequence length, making them more suitable for modeling longer metagenomic sequences. Here, we introduce Nammu, a mixed-modality Mamba-based foundation model with 167M parameters trained on the OpenMetaGenomic (OMG) corpus. Nammu is a bidirectional encoder trained with a 20K context length using a curriculum strategy, first on 64M protein sequences and then on 32M mixed-modality metagenomic contigs. We compared Nammu to gLM2, a mixed-modality transformer also trained on OMG using 37% more tokens, using taxonomy inference on a marine dataset from the Critical Assessment of Metagenome Interpretation (CAMI). Nammu outperforms gLM2 at every taxonomic level. We further assessed function via KEGG Orthology prediction in deep-sea metagenome-assembled genomes, where Nammu outperforms gLM2 (150M). These results demonstrate improved performance.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Nature Machine Intelligence
70 papers in training set
Top 0.1%
13.0%
2
Bioinformatics
1204 papers in training set
Top 2%
12.0%
3
Nature Communications
5641 papers in training set
Top 17%
11.1%
4
Nature Methods
385 papers in training set
Top 0.9%
10.7%
5
Genome Biology
637 papers in training set
Top 2%
5.5%
50% of probability mass above
6
Nature Biotechnology
172 papers in training set
Top 0.6%
5.5%
7
Cell Systems
201 papers in training set
Top 1%
4.1%
8
PLOS Computational Biology
1863 papers in training set
Top 9%
4.1%
9
BMC Bioinformatics
457 papers in training set
Top 3%
3.3%
10
Bioinformatics Advances
203 papers in training set
Top 2%
2.5%
11
Journal of Computational Biology
48 papers in training set
Top 0.5%
1.9%
12
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 26%
1.9%
13
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
1.7%
14
Nature Computational Science
55 papers in training set
Top 0.8%
1.4%
15
Scientific Reports
3612 papers in training set
Top 60%
1.4%
16
Genome Research
468 papers in training set
Top 5%
1.1%
17
Communications Biology
993 papers in training set
Top 21%
1.1%
18
Briefings in Bioinformatics
354 papers in training set
Top 6%
1.1%
19
PLOS ONE
5266 papers in training set
Top 54%
1.1%
20
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
21
Patterns
78 papers in training set
Top 2%
1.0%
22
iScience
1154 papers in training set
Top 34%
0.8%
23
Cell Reports Methods
165 papers in training set
Top 4%
0.8%
24
mSystems
394 papers in training set
Top 7%
0.6%
25
GigaScience
212 papers in training set
Top 5%
0.6%