Decoding Prokaryotic Whole Genomes with a Product-Contextualized Large Language Model
Ni, S.; Li, S.; Wang, S.; Bi, X.; Li, Y.; Gan, C.; Jin, i.; Lu, Y.; Argha, A.; Alinejad-Rokny, H.; Si, T.; Yang, M.; Wang, T.
Show abstract
Genomes encode the instructions for life, yet their full interpretation requires models capable of capturing long-range context and functional meaning at scale. Existing genome language models (gLMs) are limited by short context windows, high computational cost, and poor interpretability. We present GenSyntax, a product-contextualized large language model (LLM) trained on 49,250 annotated prokaryotic genomes. GenSyntax replaces nucleotide tokenization with gene product descriptors, transforming genomes into "genetic paragraphs" that preserve functional semantics. Using a two-stage training strategy, GenSyntax achieves leading performance in plasmid host identification, gene function prediction, genome assembly, and gene essentiality assessment compared with the other LLMs. It also enables phenotype prediction and minimal genome design, establishing a scalable and interpretable framework for genome-scale decoding and synthetic biology.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Protein Set Transformer: A protein-based genome language model to power high diversity viromics 97%
- Accuracy and data efficiency in deep learning models of protein expression 97%
- Genome-scale community modelling reveals conserved metabolic cross-feedings in epipelagic bacterioplankton communities 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.