Back

Decoding Prokaryotic Whole Genomes with a Product-Contextualized Large Language Model

Ni, S.; Li, S.; Wang, S.; Bi, X.; Li, Y.; Gan, C.; Jin, i.; Lu, Y.; Argha, A.; Alinejad-Rokny, H.; Si, T.; Yang, M.; Wang, T.

2025-12-05 genomics
10.64898/2025.12.03.692003 bioRxiv
Show abstract

Genomes encode the instructions for life, yet their full interpretation requires models capable of capturing long-range context and functional meaning at scale. Existing genome language models (gLMs) are limited by short context windows, high computational cost, and poor interpretability. We present GenSyntax, a product-contextualized large language model (LLM) trained on 49,250 annotated prokaryotic genomes. GenSyntax replaces nucleotide tokenization with gene product descriptors, transforming genomes into "genetic paragraphs" that preserve functional semantics. Using a two-stage training strategy, GenSyntax achieves leading performance in plasmid host identification, gene function prediction, genome assembly, and gene essentiality assessment compared with the other LLMs. It also enables phenotype prediction and minimal genome design, establishing a scalable and interpretable framework for genome-scale decoding and synthetic biology.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.