Back

A dual-metric framework for quantifying the biological fidelity of scRNA-seq pipelines

Su, Q.; Long, Y.; Duan, F.; Lian, Q.

2025-11-13 bioinformatics
10.1101/2025.11.12.688141 bioRxiv
Show abstract

Single-cell RNA sequencing (scRNA-seq) enables high-resolution transcriptional profiling, but early-stage processing pipelines differ markedly in barcode recovery, UMI correction, and read assignment-variations that can propagate and bias downstream analyses. We present a reproducible, parameter-aware benchmarking framework to quantify the biological fidelity of four primary pipelines-STARsolo, Cell Ranger, Kallisto|Bustools, and Alevin-fry-across simulated ground-truth datasets and complex biological contexts, including a Huntingtons disease (HD) mouse model. Our approach introduces two complementary metrics: the Cluster Annotation Score (CAS), assessing concordance between direct cell-level and cluster-level consensus labels, and the Marker Concordance Score (MCS), measuring cohesion of de novo marker genes per cell type. By systematically varying highly variable gene (n_HVG) and principal component (n_PC) settings, we map how upstream quantification interacts with downstream parameter choice. Simulations show STARsolo and Kallisto|Bustools deliver high technical accuracy, but only STARsolo consistently preserves stable cell identities and coherent marker expression across parameter regimes. In empirical datasets, alignment-based pipelines (STARsolo, Cell Ranger) yield higher CAS/MCS values and more biologically faithful annotations, while alignment-free methods show reduced signal fidelity despite faster runtimes. Differences introduced during primary processing persist after batch correction and integration, altering disease-associated cell type detection. Our open-source CAS-MCS-Scoring toolkit enables transparent evaluation of pipeline performance, providing a practical guide for selecting analysis strategies that maximize reproducibility and biological interpretability in scRNA-seq studies.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.