Back

Real-paired single-cell/bulk RNA-seq benchmark and a practical protocol for accurate cell-type deconvolution in human BAL samples

Hu, Y.; Liu, Z.; Tsao, D.; Leung, J. M.; V. Gerayeli, F.; Li, X.; Shao, X.; Sin, D.; Zhang, X.

2026-01-14 bioinformatics
10.64898/2026.01.14.699304 bioRxiv
Show abstract

BackgroundPseudo-bulk RNA-seq, generated by aggregating single-cell profiles, is widely used for benchmarking deconvolution methods because it globally approximates bulk transcriptomes and provides known cell-type proportions as ground truth. However, pseudo-bulk inherits single-cell specific measurement properties, and the extent to which these differ from real bulk RNA-seq remains difficult to quantify in the absence of paired data. In practice, deconvolution studies also commonly rely on large external single-cell references, yet the value of small, protocol-matched in-study references has not been systematically evaluated under real bulk conditions. These gaps motivate a paired benchmark that jointly examines pseudo-bulk fidelity, reference design, and their consequences for deconvolution accuracy. ResultsWe establish a fully paired benchmarking framework using split-sample, donor-matched bulk and single-cell RNA-seq (scRNA-seq) from human bronchoalveolar lavage (BAL). Embedded within a reverse five-fold cross-validation design and evaluated on real bulk RNA-seq data, this framework benchmarks 15 deconvolution algorithms published in 2013-2025 across 3 cell-type resolutions and 5 single-cell references, including three published BAL datasets, a harmonized lung BAL atlas, and an in-study reference derived from paired aliquots. We show that real bulk and matched pseudo-bulk profiles exhibit systematic gene-level differences, identifying 557 reproducibly discordant genes (|log2 FC| > 1, FDR< 0.05 by LIMMA), including cell-type informative features. These discrepancies reflect technology-specific effects and violate the linear mixing assumption underlying most deconvolution methods. We demonstrate that a protocol-matched in-study reference constructed from only six donors consistently outperforms substantially larger external references, with the advantage that increases at finer cell-type resolution. Moreover, paired samples enable the identification and selective removal of discordant genes, which further improves deconvolution accuracy for many algorithms, particularly in high-resolution settings. These findings are robust across methods, references, and evaluation criteria and extend beyond compositional accuracy to improving recovery of disease-associated cell-type differences in clinical applications. ConclusionsOur study provides the first fully paired benchmark of transcriptomic deconvolution on real bulk RNA-seq data of human BAL samples and demonstrates that reference design and data compatibility are as influential as algorithm choice. Beyond benchmarking, we introduce a practical and cost-effective protocol for deconvolution studies: generate single-cell data for a minimal subset of bulk samples (pilot pairing), use these data to construct an in-study reference and identify discordant genes, and apply the resulting insights to the full cohort. This strategy requires limited additional experimental effort yet yields substantial gains in accuracy and stability, offering actionable guidance for future bulk RNA-seq deconvolution studies across tissues and platforms. One-sentence summarySystematic benchmarking reveals pseudo-bulk biases and provides practical fixes for accurate RNA deconvolution.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.