Pre-training Genomic Language Model with Variants for Better Modeling Functional Genomics
Liu, T.; Zhang, X.; Ying, R.; Zhao, H.
Show abstract
Sequence-to-function models can predict gene expression from sequence data and be used to link genetic information with transcriptomics data to understand regulatory processes and their effects on complex phenotypes. The genomic language models are pre-trained with large-scale DNA sequences and can generate robust representations of these sequences by learning the genomic context. How-ever, few studies can estimate the predictability of gene expression levels and bridge these two classes of models together to explore individualized gene expression prediction. In this manuscript, we propose UKBioBERT as a DNA language model pre-trained with genetic variants from UK BioBank. We demonstrate that UKBioBERT generates informative embeddings capable of identifying gene functions, and improving gene expression prediction in cell lines, thereby enhancing our understanding of gene expression predictability. Building upon these embeddings, we combine UKBioBERT with state-of-the-art sequence-to-function architectures, Enformer and Borzoi, to create UKBioFormer and UKBioZoi. These models exhibit better performance in predicting highly predictable gene expression levels and can be generalized across different cohorts. Furthermore, UKBioFormer effectively captures the relationship between genetic variants and expression variations, enabling in-silico mutation analyses and eQTL identification. Collectively, our findings underscore the value of integrating genomic language models and sequence-to-function approaches for advancing functional genomics research.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning latent embedding of multi-modal single cell data and cross-modality relationship simultaneously 96%
- CMOT: Cross Modality Optimal Transport for multimodal inference 96%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 97%
- Hi-C-LSTM: Learning representations of chromatin contacts using a recurrent neural network identifies genomic drivers of conformation 96%
- DNALONGBENCH: A Benchmark Suite for Long-Range DNA Prediction Tasks 96%
Similar papers in this journal
Similar papers in this journal
- NetTIME: a multitask and base-pair resolution framework for improved transcription factor binding site prediction 97%
- Topology-Driven Negative Sampling Enhances Generalizability in Protein-Protein Interaction Prediction 96%
- DeepPHiC: Predicting promoter-centered chromatin interactions using a novel deep learning approach 96%
Similar papers in this journal
- spRefine Denoises and Imputes Spatial Transcriptomics with a Reference-Free Framework Powered by Genomic Language Model 97%
- A scalable computational framework for predicting gene expression from candidate cis-regulatory elements 96%
- CodonBERT: Large Language Models for mRNA Design and Optimization 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.