Predicting cell population-specific gene expression from genomic sequence
Michielsen, L.; Reinders, M.; Mahfouz, A.
Show abstract
Most regulatory elements, especially enhancer sequences, are cell population-specific. One could even argue that a distinct set of regulatory elements is what defines a cell population. However, discovering which non-coding regions of the DNA are essential in which context, and as a result, which genes are expressed, is a difficult task. Some computational models tackle this problem by predicting gene expression directly from the genomic sequence. These models are currently limited to predicting bulk measurements and mainly make tissue-specific predictions. Here, we present a model that leverages single-cell RNA-sequencing data to predict gene expression. We show that cell population-specific models outperform tissue-specific models, especially when the expression profile of a cell population and the corresponding tissue are dissimilar. Further, we show that our model can prioritize GWAS variants and learn motifs of transcription factor binding sites. We envision that our model can be useful for delineating cell population-specific regulatory elements.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 97%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 97%
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 96%
Similar papers in this journal
- scGPT: Towards Building a Foundation Model for Single-Cell Multi-omics Using Generative AI 98%
- Towards Universal Cell Embeddings: Integrating Single-cell RNA-seq Datasets across Species with SATURN 97%
- The Nucleotide Transformer: Building and Evaluating Robust Foundation Models for Human Genomics 97%
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 97%
- STAN, a computational framework for inferring spatially informed transcription factor activity across cellular contexts 96%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.