Benchmarking computational decontamination of ambient RNA
Cargnelli, C. B.; Nielsen, J. V.; Madsen, J.
Show abstract
Gene expression profiling of single cells using single-cell and single-nucleus RNA sequencing (sxRNA-seq) enables researchers to characterizing cellular heterogeneity and unraveling complex biological processes at unprecedented resolution. However, sxRNA-seq faces challenges due to the presence of ambient RNA, extraneous RNA molecules not originating from the cells of interest. Sample preparation is a major source of ambient RNA, where harsh conditions can lead to cell lysis and the release of intracellular RNA. This inescapable inclusion of ambient RNA can cause erroneous results and hinder downstream analyses. To address this issue, various methodologies have been developed to identify, quantify, and remove ambient RNA. Here, we rigorously evaluate 7 state-of-the-art methodologies for ambient RNA removal using simulated datasets, species-mixing experiments of varying complexities and genotype-mixing experiments. We find that no single method performs the best across all datasets and metrics, but CellBender and DecontX generally perform well.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A Universal Deep Neural Network for In-Depth Cleaning of Single-Cell RNA-Seq Data 95%
- Normalizing and denoising protein expression data from droplet-based single cell profiling 95%
- RoCK and ROI: Single-cell transcriptomics with multiplexed enrichment of selected transcripts and region-specific sequencing 95%
Similar papers in this journal
- Automated quality control and cell identification of droplet-based single-cell data using dropkick 97%
- Alignment of single-cell RNA-seq samples without over-correction using kernel density matching 95%
- Highly accurate reference and method selection for universal cross-dataset cell type annotation with CAMUS 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.