Back

DemoTape: Computational demultiplexing of targeted single-cell sequencing data

Borgsmueller, N.; Kuipers, J.; Gawron, J.; Roncador, M.; Pohly, M. F.; Acar, E.; Do, T. H. L.; Reisenauer, S. U.; Feldkamp, M. J.; Beisel, C.; Zenz, T.; Moor, A.; Beerenwinkel, N.

2024-12-10 bioinformatics
10.1101/2024.12.06.627152 bioRxiv
Show abstract

BackgroundSingle-cell sequencing can provide novel insights into the understanding and treatment of diseases. In cancer, for example, intratumor heterogeneity is a major cause of treatment resistance and relapse. Although technological progress has substantially increased the throughput of sequenced cells, single-cell sequencing remains cost and labor-intensive. Multiplexing, i.e., the pooling and subsequent joint preparation and sequencing of samples, followed by a demultiplexing step, is a common practice to reduce expenses and confounding batch effects, especially in single-cell RNA sequencing. ResultsHere, we introduce demoTape, a computational demultiplexing method for targeted single-cell DNA sequencing (scDNA-seq) data based on a distance metric between individual cells at single-nucleotide polymorphisms loci. To validate demoTape, we sequence three B-cell lymphoma patients separately and multiplexed on the Tapestri platform. We find similar genotypes, clones, and evolutionary histories in all three samples when comparing the individual with the demultiplexed samples. Using the three individually sequenced samples, we simulate multiplexed ground truth data and show that demoTape outperforms state-of-the-art demultiplexing methods designed for RNA sequencing data. Additionally, we demonstrate through downsampling that the inferred clonal composition remained largely stable for samples with fewer cells despite the inevitable loss in resolution of low-frequency clones. ConclusionsMultiplexing and subsequent genotype-based demultiplexing of scDNA-seq will reduce costs and workload, eventually allowing the sequencing of more samples. This will open new possibilities and accelerate the investigation of biological questions where cellular heterogeneity on the genomic level plays a crucial role.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.