fastdemux: Robust SNP-based demultiplexing of single-cell population genomics data
Ranjbaran, A.; Luca, F.; Pique-Regi, R.
Show abstract
Sample multiplexing reduces cost and batch effects in population based large-scale single-cell genomics studies but requires accurate and scalable computational demultiplexing. Existing genotype-based methods, such as demuxlet, provide high accuracy but can be computationally slow and memory intensive as the number of cells, donors, and informative variants increases. Here, we introduce fastdemux, a scalable genotype-based demultiplexing framework based on a diagonal linear discriminant analysis (DLDA) model that substantially improves computational efficiency while maintaining accurate donor assignment. Using a pooled single-cell RNA-seq dataset from unrelated donors, we benchmarked fastdemux against demuxlet, vireo, and demuxalot.fastdemux achieved comparable or improved demultiplexing accuracy while reducing runtime and peak memory usage by orders of magnitude relative to alternative methods. Performance remained robust across varying sequencing depths and genotype SNP filtering thresholds. In addition, the DLDA framework naturally extends to doublet and higher-order multiplet detection. We also show thatfastdemux works well with scATAC-seq data where genetic variants are more sparsely covered. Together, these results establish fastdemux as an efficient and scalable solution for genetic demultiplexing of pooled single-cell datasets.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Aberrant landscapes of maternal meiotic crossovers contribute to aneuploidies in human embryos 95%
- Automated quality control and cell identification of droplet-based single-cell data using dropkick 95%
- Alignment of single-cell RNA-seq samples without over-correction using kernel density matching 94%
Similar papers in this journal
- scConsensus: combining supervised and unsupervised clustering for cell type identification in single-cell RNA sequencing data 94%
- nPoRe: n-Polymer Realigner for improved pileup variant calling 94%
- eSVD-DE: Cohort-wide differential expression in single-cell RNA-seq data using exponential-family embeddings 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.