Back

scDblFinder in Python with GPU support

Hiropedi, A.; Germain, P.-L.

2026-08-20 bioinformatics
10.64898/2026.08.12.744148 bioRxiv
Show abstract

High-throughput single-cell sequencing provides a scalable solution for characterizing cells and profiling gene expression for hundreds to millions of cells. However, this process gives rise to doublets, which can lead to inaccurate conclusions drawn from the data. A number of packages have therefore been developed to help accurately detect them, and in particular scDblFinder has been shown to outperform alternatives in the detection of doublets in single-cell (RNA) sequencing data. Being implemented in R, however, its adoption has been more limited in the Python community. Here, we present scDblFinderPy, a Python-based implementation of the scDblFinder R method, and show that it obtains similar performances. Furthermore, we include in it optional GPU support, thus further speeding up the process.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.