Canonical self-supervised pretraining paradigm constrains the capacity of genomic language models on regulatory decoding
Liang, Y.-X.; Wang, Y.; Pan, W.-Y.; Chen, Z.-Y.; Wei, J.-C.; Gao, G.
Show abstract
Recent studies suggest that genomic language models (gLMs) could help decode genomic regulatory code. Here, we systematically evaluated 11 representative gLMs across multiple regulatory genomics applications and found that current gLMs offer limited advantages over the random baseline. Further analysis revealed a systematic misalignment between the canonical sequence-only self-supervised pretraining paradigm and the context-specific dynamic nature of gene regulation, highlighting the need for function-oriented pretraining strategies that explicitly incorporate biochemical and regulatory priors.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Landscape of allele-specific transcription factor binding in the human genome 97%
- Predicting cell type-specific epigenomic profiles accounting for distal genetic effects 97%
- G4mer: An RNA language model for transcriptome-wide identification of G-quadruplexes and disease variants from population-scale genetic data 97%
Similar papers in this journal
- Deep-learning prediction of gene expression from personal genomes 97%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 97%
- Evaluating the representational power of pre-trained DNA language models for regulatory genomics 97%
Similar papers in this journal
- Integrating convolution and self-attention improves language model of human genome for interpreting non-coding regions at base-resolution 97%
- Predicting gene expression from histone marks using chromatin deep learning models depends on histone mark function, regulatory distance and cellular states 97%
- DeepCLIP: Predicting the effect of mutations on protein-RNA binding with Deep Learning 96%
Similar papers in this journal
- Quantitative occupancy of myriad transcription factors from one DNase experiment enables efficient comparisons across conditions 97%
- Prioritization of enhancer mutations by combining allele-specific chromatin accessibility with deep learning 97%
- DeepArk: modeling cis-regulatory codes of model species with deep learning 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.