Guided tokenization and domain knowledge enhance genomic language models' performance
Mahangade, V.; Mollerus, M.; Crandall, K. A.; Rahnavard, A.
Show abstract
Adapting language models to genomic and metagenomic sequences presents unique challenges, particularly in tokenization and task-specific generalization. Standard methods, such as fixed-length k-mers or byte pair encoding, often fail to preserve biologically meaningful patterns essential for downstream tasks. We introduce Guided Tokenization (GT), a strategy that prioritizes biologically and statistically important subsequences based on importance scores, model attention, and class distributions. Combined with domain adaptation, which incorporates prior domain specific biological knowledge, this approach improves both representation quality and classification accuracy in compact genomic language models (gLMs). GT enhances biological awareness in genomic language models, particularly for effective small and mid-sized models across key tasks, including DNA sequence read classification, promoter detection, antimicrobial resistance classification, and targeted amplicon taxonomic profiling. Our results highlight the promise of guided tokenization and domain-aware modeling for building efficient, biologically grounded language models for scalable genomic applications.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Protein Set Transformer: A protein-based genome language model to power high diversity viromics 95%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 94%
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 94%
Similar papers in this journal
- A k-mer-based maximum likelihood method for estimating distances of reads to genomes enables genome-wide phylogenetic placement. 96%
- DEMINERS enables clinical metagenomics and comparative transcriptomic analysis by increasing throughput and accuracy of nanopore direct RNA sequencing 94%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 94%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 94%
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 94%
- Unraveling the start element and regulatory divergence of core promoters across the domain Bacteria 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.