Multiplets in scRNA-seq data: extent of the problem and efficacy of methods for removal
Ttoouli, D.; Hoffmann, D.
Show abstract
Multiplets--droplets that capture more than one cell--are a known artefact in droplet-based single-cell RNA sequencing (scRNA-seq), yet their prevalence and impact remain underestimated. In this study, we assess the frequency of multiplets across diverse publicly available datasets and evaluate how well commonly used detection tools are able to identify them. Using cell hashing data to determine a lower bound of the true multiplet rate, we demonstrate that commonly used heuristic estimations systematically underestimate multiplet rates, and that existing tools--despite optimized parameters--detect only a small subset of cell-hashing multiplets. We further refine a Poisson-based model to estimate the true multiplet rate, revealing that actual rates can exceed heuristic predictions by more than twofold. Downstream analyses are significantly affected by multiplets: they are not confined to isolated clusters but are distributed throughout the transcriptional landscape, where they distort clustering and cell type annotation. Using both quantitative and qualitative approaches, we visualize these effects and show that cell-hashing-informed multiplet removal eliminates artefactual clusters and improves annotation clarity, whereas computationally detected multiplets fail to fully remove artefacts in the most common experimental contexts. Our findings confirm that multiplet contamination remains a pervasive and under-addressed issue in scRNA-seq analysis. Since most datasets lack multiplexing, researchers must often rely on heuristics and limited tools, leaving many multiplets unidentified. We advocate for more robust multiplet-detection strategies, including multimodal validation, to ensure more accurate and interpretable scRNA-seq results.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 95%
- STREAK: A Supervised Cell Surface Receptor Abundance Estimation Strategy for Single Cell RNA-Sequencing Data using Feature Selection and Thresholded Gene Set Scoring 95%
- HiCImpute: A Bayesian Hierarchical Model for Identifying Structural Zeros and Enhancing Single Cell Hi-C Data. 95%
Similar papers in this journal
- scDeepInsight: a supervised cell-type identification method for scRNA-seq data with deep learning 95%
- Comparison of High-Throughput Single-Cell RNA Sequencing Data Processing Pipelines 95%
- How does data structure impact cell-cell similarity? Evaluating the influence of structural properties on proximity metric performance in single cell RNA-seq data 95%
Similar papers in this journal
- Probability of stealth multiplets in sample-multiplexing for droplet-based single-cell analysis 96%
- Differential detection workflows for multi-sample single-cell RNA-seq data 94%
- FastCAR: Fast Correction for Ambient RNA to facilitate differential gene expression analysis in single-cell RNA-sequencing datasets 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.