Impacts of Cell Ranger versions on Chromium gene expression data
Abugessaisa, I.; Hasegawa, A.; Walker, S.; Katayama, S.; Kere, J.; Kasukawa, T.
Show abstract
In droplet-based single cell gene expression data, cell barcode processing by Cell Ranger (CR) is a standard pipeline. But no systematic evaluation of the impact of CR version on single cell gene expression data has been conducted. To comprehensively evaluate the impact of CR version, we considered six molecular quality criteria, quantified gene expression, and performed downstream analysis for 12 single-cell datasets. Each dataset was processed by 15 versions of CR. We demonstrated that different versions of CR yield different numbers of cell barcodes with significant variation in detected UMIs, features, molecular qualities and average gene expression of protein-coding and lncRNA for the same dataset. Our analysis finds distinction between two diverse categories of cell barcodes: common barcodes unmasked by all versions of CR, and specific barcodes only unmasked/masked by some versions. Surprisingly, we observed variation in molecular read-out between common cell barcodes when called by different versions of CR. The specific barcodes yield skewed gene body coverage and form distinct clusters. The choice of CR version affects scores for quality, average gene expression, clustering results, and top cluster marker genes of the dataset.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scROSHI - robust supervised hierarchical identification of single cells 95%
- Fast analysis of Spatial Transcriptomics (FaST): an ultra lightweight and fast pipeline for the analysis of high resolution spatial transcriptomics. 95%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 95%
Similar papers in this journal
- Revealing the Prevalence of Suboptimal Cells and Organs in Reference Cell Atlases: An Imperative for Enhanced Quality Control 96%
- Illuminating the dark side of the human transcriptome with TAMA Iso-Seq analysis 95%
- Choice of pre-processing pipeline influences clustering quality of scRNA-seq datasets 95%
Similar papers in this journal
- QClus: A droplet-filtering algorithm for enhanced snRNA-seq data quality in challenging samples 96%
- MarcoPolo: a clustering-free approach to the exploration of differentially expressed genes along with group information in single-cell RNA-seq data 96%
- LINE-1 Retrotransposon expression in cancerous, epithelial and neuronal cells revealed by 5'-single cell RNA-Seq 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.