scDblFinder in Python with GPU support
Hiropedi, A.; Germain, P.-L.
Show abstract
High-throughput single-cell sequencing provides a scalable solution for characterizing cells and profiling gene expression for hundreds to millions of cells. However, this process gives rise to doublets, which can lead to inaccurate conclusions drawn from the data. A number of packages have therefore been developed to help accurately detect them, and in particular scDblFinder has been shown to outperform alternatives in the detection of doublets in single-cell (RNA) sequencing data. Being implemented in R, however, its adoption has been more limited in the Python community. Here, we present scDblFinderPy, a Python-based implementation of the scDblFinder R method, and show that it obtains similar performances. Furthermore, we include in it optional GPU support, thus further speeding up the process.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- DeepImpute: an accurate, fast and scalable deep neural network method to impute single-cell RNA-Seq data 96%
- A comparison of marker gene selection methods for single-cell RNA sequencing data 94%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 94%
Similar papers in this journal
- pyALRA: python implementation of low-rank zero-preserving approximation of single cell RNA-seq 95%
- ScaleSC: A superfast and scalable single cell RNA-seq data analysis pipeline powered by GPU. 95%
- kmtricks: Efficient and flexible construction of Bloom filters for large sequencing data collections 94%
Similar papers in this journal
- Fast analysis of Spatial Transcriptomics (FaST): an ultra lightweight and fast pipeline for the analysis of high resolution spatial transcriptomics. 94%
- MUFFIN : A suite of tools for the analysis of functional sequencing data 94%
- SIQ: easy quantitative measurement of mutation profiles in sequencing data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.