Assessment of dispersion metrics for estimating single-cell transcriptional variability
Chen, T. T.; Boyer, L. A.; Agarwal, D.
Show abstract
Single-cell RNA sequencing data enables analysis of transcript levels of single cells across different cell types and conditions. Recent work has highlighted the value of measuring gene-specific transcriptional variability, or noise, within a genetically identical population of cells in addition to mean expression given that these differences contribute to biological processes including development and disease. However, measuring transcriptional noise remains a challenge. Here, we systematically compared statistical methods by simulating single-cell data by varying both dispersion and count size to assess the relative responsiveness to noise of several commonly used statistical metrics: the Gini index, variance-to-mean ratio, variance, and Shannon entropy. We found that the variance-to-mean ratio scales approximately linearly with increasing dispersion and is scale-invariant. In contrast, the Gini index displayed paradoxical behavior, and Shannon entropy was not scale-invariant. Thus, we next applied the variance-to-mean ratio to measure transcriptional variability in a publicly available single-cell dataset of embryonic hearts from a mouse model of maternal hyperglycemia. Our data show that many genes display transcriptional variability within the same cell type, and that this variation does not correlate with gene characteristics such as transcript level, promoter GC content, or evolutionary gene age. Notably, many of the genes and pathways with highest transcriptional variability were not identified as differentially expressed, and have in fact been implicated in maternal hyperglycemia in other studies, suggesting that transcriptional variability can provide additional biologically relevant information beyond what is observed from studying mean expression alone.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GoM DE: interpreting structure in sequence count data with differential expression analysis allowing for grades of membership 95%
- Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics 94%
- Modelling group heteroscedasticity in single-cellRNA-seq pseudo-bulk data 94%
Similar papers in this journal
Similar papers in this journal
- Probability of stealth multiplets in sample-multiplexing for droplet-based single-cell analysis 93%
- Differential detection workflows for multi-sample single-cell RNA-seq data 93%
- Functional module detection through integration of single-cell RNA sequencing data with protein-protein interaction networks 93%
Similar papers in this journal
- Tissue specificity-aware TWAS (TSA-TWAS) framework identifies novel associations with metabolic, immunologic, and virologic traits in HIV-positive adults 94%
- Meta-imputation of transcriptome from genotypes across multiple datasets using summary-level data. 94%
- Mouse-Geneformer: A Deep Learning Model for Mouse Single-Cell Transcriptome and Its Cross-Species Utility 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.