CMDdemux: an efficient single cell demultiplexing method
Wang, J.; Chen, L.; Brown, D. V.; Chiu, C.; Speed, T. P.
Show abstract
Multiplexing technologies label cells with molecular tags, allowing cells from different donors to be pooled together for sequencing. Although this approach enhances cell throughput, eliminates batch effects, and enables doublet detection, limitations of hashtag-based labelling can still lead to low-quality data. Existing demultiplexing methods can accurately assign donor identities in high-quality datasets, but they often fail on low-quality data. To address this, we developed CMDdemux, a method comprising three key steps: within-cell centered log-ratio (CLR) normalization of hashtag count data, K-medoids clustering, and classification of cells based on Mahalanobis distance. By integrating both hashing and mRNA data, CMDdemux achieves high accuracy in distinguishing singlets, doublets, and negatives. It also provides visualization tools to help users inspect potentially misclassified droplets. We benchmarked CMDdemux against existing methods using a range of high- and low-quality datasets. Results show that CMDdemux consistently outperforms other approaches, demonstrating robust performance on both high- and low-quality data where other methods fail. CMDdemux is particularly effective in handling diverse types of low-quality multiplexing data across different multiplexing technologies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 97%
- scSemiProfiler: Advancing Large-scale Single-cell Studiesthrough Semi-profiling with Deep Generative Models andActive Learning 97%
- uniPort: a unified computational framework for single-cell data integration with optimal transport 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.