OddSNP: a predictive framework for optimizing multiplexed single-cell RNA-seq experiments
Allendes Osorio, R. S.; Nishimura, T.; Shigihara, Y.; Takebe, T.; Nemoto, T.
Show abstract
Donor multiplexing is a powerful strategy to increase scale, lower the costs, and reduce batch effects in single-cell RNA sequencing (scRNAseq), but clear guidelines for experimental design are lacking, forcing researchers to risk costly demultiplexing failures. To address this, we introduce SNP-Information Content (SNP-IC), a quantitative metric computable from simple unpooled pilot data that accurately predicts the success of genotype-based demultiplexing. Across multiple large-scale datasets using stem cell and organoid models, we establish a robust SNP-IC threshold of approximately 50, above which cells can be reliably assigned to their donor of origin. For more challenging genotype-free approaches, we define a pairwise metric, cpSNP-IC, and demonstrate a much higher requirement of approximately 3,000. Our open-source framework, oddSNP, implements this predictive model, allowing researchers to perform in silico titrations of sequencing depth and donor complexity to optimize experimental design before committing to large-scale studies. oddSNP provides a practical framework, enabling researchers to strategically optimize sequencing depth and donor numbers to maximize experimental success while managing costs and minimizing the risk of catastrophic data loss.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CelLink: integrating single-cell multi-omics data with weak feature linkage and imbalanced cell populations 95%
- MarcoPolo: a clustering-free approach to the exploration of differentially expressed genes along with group information in single-cell RNA-seq data 94%
- QClus: A droplet-filtering algorithm for enhanced snRNA-seq data quality in challenging samples 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.