Back

Error Correction Algorithms for Efficient Gene ExpressionQuantification in Single Cell Transcriptomics

Zentgraf, J.; Schmitz, J. E.; Keller, A.; Rahmann, S.

2025-12-01 bioinformatics
10.1101/2025.11.27.690682 bioRxiv
Show abstract

Technological advances in single-cell RNA sequencing (scRNA-seq) allow us to sequence the transcriptomes of thousands of single cells in parallel, resulting in massive amounts of raw sequence data that must be processed efficiently to obtain a genes x cells expression matrix. In droplet-based scRNA-seq protocols, the sequenced mRNA molecules are tagged with a cell-specific barcode and a unique molecular identifier (UMI) within each cell. Both barcodes and UMIs may contain errors from production, amplification, or sequencing. Correcting and resolving such errors before further processing yields more reliable data and more accurate expression measurements. We propose algorithmic advancements for barcode correction, read-to-gene mapping, and UMI resolution, which we combine into a new method called O_SCPLOWARCANEC_SCPLOW for efficient gene expression quantification from scRNA-seq data. We additionally provide an implementation as a workflow-friendly command-line tool, also called O_SCPLOWARCANEC_SCPLOW. This work builds on the recently published Fourway method to efficiently discover DNA k-mers with a Hamming distance of 1, speeding up barcode correction and UMI resolution and allowing for distinguishing k-mers into weakly and strongly unique ones during read-to-gene mapping. As a side result of separate interest, we show that for the mapping step, it suffices to store three genes per k-mer in order to cover almost all of the genes almost completely, thus avoiding arbitrarily large color sets in the colored De Bruijn graph index. As a result, O_SCPLOWARCANEC_SCPLOW is faster than existing methods while producing very similar results, as demonstrated in a comparison with CO_SCPLOWELLC_SCPLOWRO_SCPLOWANGERC_SCPLOW, KO_SCPLOWALLISTOC_SCPLOWO_SCPCAP|C_SCPCAPO_SCPLOWBUSTOOLSC_SCPLOW and AO_SCPLOWLEVINC_SCPLOWO_SCPCAP-C_SCPCAPO_SCPLOWFRYC_SCPLOW. O_SCPLOWARCANEC_SCPLOW is available via GitLab (https://gitlab.com/rahmannlab/arcane).

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.