Genomic Touchstone: Benchmarking Genomic Language Models in the Context of the Central Dogma
Wang, Y.; Cai, Z.; Zeng, Q.; Gao, Y.; Ouyang, J.; Xu, Y.; Yang, S.; He, S.; Nie, Y.; Cai, Y.; Zhou, F.; Jin, C.; Wang, X.; Xie, Z.; Zhu, D.; Xie, T.; Cheng, K.-T.; Yang, C.; Fu, X.; Wang, J.; Zhang, K.; Yao, J.; Rabadan, R.; Chen, H.
Show abstract
The emergence of genomic language models (gLMs) has revolutionized the analysis of genomic sequences, enabling robust capture of biologically meaningful patterns from DNA sequences for an improved understanding of human genome-wide regulatory programs, variant pathogenicity and therapeutic discovery. Given that DNA serves as the foundational blueprint within the central dogma, the ultimate evaluation of a gLM is its ability to generalize across this entire biological cascade. However, existing evaluations lack this holistic and consistent framework, leaving researchers uncertain about which model best translates sequence understanding into downstream biological prediction. Here we present Genomic Touchstone, a comprehensive benchmark designed to evaluate gLMs across 36 diverse tasks and 88 datasets structured along the central dogmas modalities of DNA, RNA, and protein, encompassing 5.34 billion base pairs of genomic sequences. We evaluate 34 representative models encompassing transformers and convolutional neural networks, as well as emerging efficient architectures such as Hyena and Mamba. Our analysis yield four key insights. First, gLMs achieve comparable or superior performance on RNA and protein tasks compared to models pretrained on these molecules. Second, transformer-based models continue to lead in overall performance, yet efficient sequence models show promising task-specific capabilities and deserve further exploration. Third, the scaling behavior of gLMs remains incompletely understood. While longer input sequences and more diverse pretraining data consistently improve performance, increases in model size do not always translate into better results. Fourth, pretraining strategies, including the choice of training objectives and the composition of pretraining corpora, exert substantial influence on downstream generalization across different genomic contexts. Genomic Touchstone establishes a unified evaluation framework tailored to human genomics. By spanning multiple molecular modalities and biological tasks, it offers a valuable foundation for guiding future gLM design and understanding model generalizability in complex biological contexts.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking Pre-trained Genomic Language Models for RNA Sequence-Related Predictive Applications 98%
- CodonTransformer: a multispecies codon optimizer using context-aware neural networks 97%
- GRouNdGAN: GRN-guided simulation of single-cell RNA-seq data using causal generative adversarial networks 96%
Similar papers in this journal
- Correcting gradient-based interpretations of deep neural networks for genomics 97%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 97%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 97%
Similar papers in this journal
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 95%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 95%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.