Back

SingleRust: A High-Performance Toolkit for Single-Cell Data Analysis at Scale

Diks, I. F.; Flotho, M.; Keller, A.

2025-08-05 bioinformatics
10.1101/2025.08.04.668429 bioRxiv
Show abstract

Single-cell RNA sequencing studies increasingly generate datasets exceeding 10 million cells, surpassing the memory capacity of standard analytical tools on typical institutional infrastructure. Here we introduce SingleRust, a computational framework that addresses these constraints through systematic algorithmic optimizations and systems-level design. Key improvements include sparse masked principal component analysis that reduces memory footprint while preserving biological signal, lock-free parallel implementations for differential expression testing, and adaptive k-nearest neighbor algorithms that automatically select optimal data structures based on dataset size. These optimizations achieve 2.4-25.5-fold performance improvements and 1.3-3.0-fold memory reduction compared to Scanpy, enabling routine analysis of 30 million cells on our representative test system with 512 GB RAM. Comprehensive validation confirms numerical equivalence with established methods while maintaining biological interpretation fidelity. SingleRust maintains full compatibility with the AnnData ecosystem while providing researchers immediate access to population-scale analyses on existing infrastructure, addressing a critical bottleneck in single-cell genomics workflows.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.