Benchmarking large-scale single-cell RNA-seq analysis
Billato, I.; Pages, H.; Carey, V.; Waldron, L.; Sales, G.; Romualdi, C.; Risso, D.
Show abstract
The increasing size of single-cell RNA sequencing (scRNA-seq) datasets poses major computational challenges. This work benchmarks the scalability, efficiency, and accuracy of five widely used analysis frameworks (Seurat, OSCA, scrap-per, Scanpy, and rapids singlecell), focusing on the impact of algorithmic and infrastructural choices on performance. We performed a systematic comparison of these workflows using representative datasets, including a 1.3 million mouse brain cell dataset for scalability and three smaller datasets (BE1, scMixology, and cord blood CITE-seq) with ground truth labels to assess clustering accuracy. Principal Component Analysis (PCA) was used as a paradigmatic step to evaluate the computational performance of six SVD algorithms (exact, ARPACK, IRLBA, randomized, Jacobi, and incremental PCA) across multiple data representations (dense, sparse, HDF5) and hardware configurations (CPU vs GPU). All methods showed high concordance in PCA results, with negligible loss of accuracy in truncated approaches. GPU-based computation using rapids singlecell provided a 15x speed-up over the best CPU methods, with moderate memory usage. On CPU, ARPACK and IRLBA were the most efficient for sparse matrices, while randomized SVD performed best for HDF5-backed data. Among full pipelines, rapids singlecell was the fastest, whereas OSCA and scrapper achieved the highest clustering accuracy (ARI up to 0.97) in datasets with known cell identities. Performance differences were largely driven by the choice of highly variable genes (HVGs) and PCA implementation. The study highlights that scalability in scRNA-seq analysis depends critically on both algorithmic and infrastructural factors. GPU acceleration and optimized BLAS/LAPACK configurations markedly enhance performance, while Bioconductor-based pipelines remain robust in accuracy. The provided benchmarks offer practical guidelines for efficient and reliable analysis of large-scale single-cell datasets.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A comparison of marker gene selection methods for single-cell RNA sequencing data 97%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 97%
- A systematic evaluation of highly variable gene selection methods for single-cell RNA-sequencing 96%
Similar papers in this journal
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 96%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 96%
- HiCImpute: A Bayesian Hierarchical Model for Identifying Structural Zeros and Enhancing Single Cell Hi-C Data. 96%
Similar papers in this journal
- SCEMENT: Scalable and Memory Efficient Integration of Large-scale Single Cell RNA-sequencing Data 96%
- CCC-GPU: A graphics processing unit (GPU)-accelerated nonlinear correlation coefficient for large-scale transcriptomic analyses 96%
- Adjustment of spurious correlations in co-expression measurements from RNA-Sequencing data 95%
Similar papers in this journal
- Variance-adjusted Mahalanobis (VAM): a fast and accurate method for cell-specific gene set scoring 96%
- edgeR v4: powerful differential analysis of sequencing data with expanded functionality and improved support for small counts and larger datasets 96%
- Assessing the impact of transcriptomics data analysis pipelines on downstream functional enrichment results 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.