PatchDNA: A Flexible and Biologically-Informed Alternative to Tokenization for DNA
Del Vecchio, A.; Kapourani, C.-A.; Athar, A. M.; Dobrowolska, A.; Anighoro, A.; Tenmann, B.; Edwards, L.; Regep, C.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWDNA language models are emerging as powerful tools for representing genomic sequences, with recent progress driven by self-supervised learning. However, performance on downstream tasks is sensitive to tokenization strategies reflecting the complex encodings in DNA, where both regulatory elements and single-nucleotide changes can be functionally significant. Yet existing models are fixed to their initial tokenization strategy; single-nucleotide encodings result in long sequences that challenge transformer architectures, while fixed multi-nucleotide schemes like byte pair encoding struggle with character level modeling. Drawing inspiration from the Byte Latent Transformers combining of bytes into patches, we propose that patching provides a competitive and more efficient alternative to tokenization for DNA sequences. Furthermore, patching eliminates the need for a fixed vocabulary, which offers unique advantages to DNA. Leveraging this, we propose a biologically informed strategy, using evolutionary conservation scores as a guide for patch boundaries. By prioritizing conserved regions, our approach directs computational resources to the most functionally relevant parts of the DNA sequence. We show that models up to an order of magnitude smaller surpass current state-of-the-art performance in existing DNA benchmarks. Importantly, our approach provides the flexibility to change patching without retraining, overcoming a fundamental limitation of current tokenization methods.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 95%
- Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data 95%
- Pair consensus decoding improves accuracy of neural network basecallers for nanopore sequencing 94%
Similar papers in this journal
- Learning probabilistic protein-DNA recognition codes from DNA-binding specificities using structural mappings 95%
- Assessing transcriptomic re-identification risks using discriminative sequence models 95%
- Joint imputation and deconvolution of gene expression across spatial transcriptomics platforms 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.