From nucleotides to semantics: genomic representation learning via joint-embedding predictive architecture
Wang, C.; Qi, Q.; Sun, H.; Zhuang, Z.; He, B.; Liu, S.; Liao, J.; Wang, J.
Show abstract
Decoding the regulatory syntax encoded in genomic sequences is a central objective in computational biology. Most existing genomic foundation models treat DNA as a language and adopt pretraining objectives from natural language processing. DNA sequences, however, lack explicit semantic boundaries and contain substantial evolutionary noise. Nucleotide-level reconstruction in a low-dimensional input space can therefore increase computational overhead and may yield representations with limited discriminative capacity. Downstream tasks often depend on expensive finetuning, which restricts practical use in many biology laboratories. Here we present GenoJEPA, a genomic representation learning framework based on joint-embedding predictive architecture. GenoJEPA combines continuous patching with semantic alignment, shifting the optimization from local base reconstruction to semantic alignment in latent space. Across 55 downstream tasks, GenoJEPA shows strong representational capacity and robust generalization while reducing parameter count and computational cost. The resulting semantic vectors from frozen GenoJEPA support lightweight GPU-free classifiers to achieve competitive accuracy. These results suggest a practical route towards efficient training and broad application of larger-scale genomic foundation models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 96%
- Sampling from Disentangled Representations of Single-Cell Data Using Generative Adversarial Networks 96%
- Pair consensus decoding improves accuracy of neural network basecallers for nanopore sequencing 96%
Similar papers in this journal
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 96%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 96%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.