Fine-grained deconvolution of cell-type effects from human bulk brain data using a large single-nucleus RNA sequencing based reference panel.
van den Oord, E. J.; Aberg, K. A.
Show abstract
Brain disorders are leading causes of disability worldwide. Gene expression studies provide promising opportunities to better understand their etiology. When studying bulk tissue, cellular diversity may cause many genes that are differentially expressed in cases and controls to remain undetected. Furthermore, identifying the specific cell-types from which association signals originate is key to formulating refined hypotheses of disease etiology, designing proper follow-up experiments and, eventually, developing novel clinical interventions. Cell-type effects can be deconvoluted statistically from bulk expression data using cell-type proportions estimated with the help of a reference panel. To create a fine-grained reference panel for the human prefrontal cortex, we analyzed data from the seven largest single nucleus RNA-seq (snRNA-seq) studies. Seventeen cell-types were robustly detected across all seven studies. To estimate the cell-type proportions, we proposed an empirical Bayes estimator that is suitable for the new panel that involves multiple low abundant cell-types. Furthermore, to avoid the use of a very large reference panel and prevent challenges with public access of nuclei level data, our estimator uses a panel comprising mean expression levels rather than the nuclei level snRNA-seq data. Evaluations show that our empirical Bayes estimator produces highly accurate and unbiased cell-type proportion estimates. Transcriptome-wide association studies performed with permuted bulk RNA-seq data showed that it is possible to perform TWASs for even the rarest cell-types without an increased risk of false positives. Furthermore, we determined that for optimal statistical power the best approach is to analyze all cell-types in the panel as opposed to grouping or omitting (rare) cell-types.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Characterization of the nuclear and cytosolic transcriptomes in human brain tissue reveals new insights into the subcellular distribution of RNA transcripts 95%
- Altered gene expression profiles impair the nervous system development in individuals with 15q13.3 microdeletion 94%
- Transcriptome-wide high-throughput mapping of protein-RNA occupancy profiles using POP-seq 93%
Similar papers in this journal
- Comprehensive evaluation of human brain gene expression deconvolution methods 96%
- Human-lineage-specific genomic elements: relevance to neurodegenerative disease and APOE transcript usage 95%
- Identifying cell type specific driver genes in autism-associated copy number loci from cerebral organoids 95%
Similar papers in this journal
- Discovery of disease-associated cellular states using ResidPCA in single-cell RNA and ATAC sequencing data 93%
- Powerful eQTL mapping through low coverage RNA sequencing 93%
- Scalable Bayesian functional GWAS method accounting for multivariate quantitative functional annotations with applications to studying Alzheimer’s disease 92%
Similar papers in this journal
Similar papers in this journal
- Data-driven Identification of Total RNA Expression Genes (TREGs) for Estimation of RNA Abundance in Heterogeneous Cell Types 95%
- Benchmark of cellular deconvolution methods using a multi-assay reference dataset from postmortem human prefrontal cortex 95%
- Cell-type specific inference from bulk RNA-sequencing data by integrating single cell reference profiles via EPIC-unmix 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.