The effect of background noise and its removal on the analysis of single-cell expression data
Janssen, P.; Kliesmete, Z.; Vieth, B.; Adiconis, X.; Simmons, S.; Marshall, J.; McCabe, C.; Heyn, H.; Levin, J. Z.; Enard, W.; Hellmann, I.
Show abstract
BACKGROUNDIn droplet-based single-cell and single-nucleus RNA-seq experiments, not all reads associated with one cell barcode originate from the encapsulated cell. Such background noise is attributed to spillage from cell-free ambient RNA or barcode swapping events. Here, we characterize this background noise exemplified by three single-cell RNA-seq (scRNA-seq) and two single-nucleus RNA-seq (snRNA-seq) replicates of mouse kidney cells. For each experiment, kidney cells from two mouse subspecies were pooled, allowing to identify cross-genotype contaminating molecules and estimate the levels of background noise. RESULTSWe find that background noise is highly variable across replicates and individual cells, making up on average 3-35% of the total counts (UMIs) per cell and show that this has a considerable impact on the specificity and detectability of marker genes. In search of the source of background noise, we find that expression profiles of cell-free droplets are very similar to expression profiles of cross-genotype contamination and hence that the majority of background molecules originates from ambient RNA. Finally, we use our genotype-based estimates to evaluate the performance of three methods (CellBender, DecontX, SoupX) that are designed to quantify and remove background noise. We find that CellBender provides the most precise estimates of background noise levels and also yields the highest improvement for marker gene detection. By contrast, clustering and classification of cells are fairly robust towards background noise and only small improvements can be achieved by background removal that may come at the cost of distortions in fine structure. CONCLUSIONOur findings help to better understand the extent, sources and impact of background noise in single-cell experiments and provide guidance on how to deal with it.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of cell barcodes from long-read single-cell RNA-seq with BLAZE 96%
- scCDC: a computational method for gene-specific contamination detection and correction in single-cell and single-nucleus RNA-seq data 96%
- Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics 96%
Similar papers in this journal
- On the discovery of population-specific state transitions from multi-sample multi-condition single-cell RNA sequencing data 96%
- Atlas-scale single-cell multi-sample multi-condition data integration using scMerge2 96%
- Normalisr: normalization and association testing for single-cell CRISPR screen and co-expression 96%
Similar papers in this journal
- Inferring cell diversity in single cell data using consortium-scale epigenetic data as a biological anchor for cell identity 95%
- Flexible comparison of batch correction methods for single-cell RNA-seq using BatchBench 95%
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 95%
Similar papers in this journal
- Single-cell Mayo Map (scMayoMap): an easy-to-use tool for cell type annotation in single-cell RNA-sequencing data analysis 94%
- Single-cell DNA and RNA sequencing reveals the dynamics of intra-tumor heterogeneity in a colorectal cancer model 93%
- Differentiation is accompanied by a progressive loss in transcriptional memory 93%
Similar papers in this journal
- Probability of stealth multiplets in sample-multiplexing for droplet-based single-cell analysis 95%
- Revealing the Prevalence of Suboptimal Cells and Organs in Reference Cell Atlases: An Imperative for Enhanced Quality Control 94%
- MitoDelta: identifying mitochondrial DNA deletions at cell-type resolution from single-cell RNA sequencing data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.