Decoding sequence determinants of gene expression in diverse cellular and disease states
Lal, A.; Karollus, A.; Gunsalus, L.; Garfield, D.; Nair, S.; Tseng, A. M.; Gordon, M. G.; Collier, J. L.; Diamant, N.; Biancalani, T.; Corrada Bravo, H.; Scalia, G.; Eraslan, G.
Show abstract
Sequence-to-function models that predict gene expression from genomic DNA sequence have proven valuable for many biological tasks, including understanding cis-regulatory syntax and interpreting non-coding genetic variants. However, current state-of-the-art models have been trained largely on bulk expression profiles from healthy tissues or cell lines, and have not learned the properties of precise cell types and states that are captured in large-scale single-cell transcriptomic datasets. Thus, they lack the ability to perform these tasks at the resolution of specific cell types or states across diverse tissue and disease contexts. To address this gap, we present Decima, a model that predicts the cell type- and condition- specific expression of a gene from its surrounding DNA sequence. Decima is trained on single-cell or single-nucleus RNA sequencing data from over 22 million cells, and successfully predicts the cell type-specific expression of unseen genes based on their sequence alone. Here, we demonstrate Decimas ability to reveal the cis-regulatory mechanisms driving cell type-specific gene expression and its changes in disease, to predict non-coding variant effects at cell type resolution, and to design regulatory DNA elements with precisely tuned, context-specific functions.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leveraging supervised learning for functionally-informed fine-mapping of cis-eQTLs identifies an additional 20,913 putative causal eQTLs 96%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
- The molecular basis, genetic control and pleiotropic effects of local gene co-expression 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.