Real-paired single-cell/bulk RNA-seq benchmark and a practical protocol for accurate cell-type deconvolution in human BAL samples
Hu, Y.; Liu, Z.; Tsao, D.; Leung, J. M.; V. Gerayeli, F.; Li, X.; Shao, X.; Sin, D.; Zhang, X.
Show abstract
BackgroundPseudo-bulk RNA-seq, generated by aggregating single-cell profiles, is widely used for benchmarking deconvolution methods because it globally approximates bulk transcriptomes and provides known cell-type proportions as ground truth. However, pseudo-bulk inherits single-cell specific measurement properties, and the extent to which these differ from real bulk RNA-seq remains difficult to quantify in the absence of paired data. In practice, deconvolution studies also commonly rely on large external single-cell references, yet the value of small, protocol-matched in-study references has not been systematically evaluated under real bulk conditions. These gaps motivate a paired benchmark that jointly examines pseudo-bulk fidelity, reference design, and their consequences for deconvolution accuracy. ResultsWe establish a fully paired benchmarking framework using split-sample, donor-matched bulk and single-cell RNA-seq (scRNA-seq) from human bronchoalveolar lavage (BAL). Embedded within a reverse five-fold cross-validation design and evaluated on real bulk RNA-seq data, this framework benchmarks 15 deconvolution algorithms published in 2013-2025 across 3 cell-type resolutions and 5 single-cell references, including three published BAL datasets, a harmonized lung BAL atlas, and an in-study reference derived from paired aliquots. We show that real bulk and matched pseudo-bulk profiles exhibit systematic gene-level differences, identifying 557 reproducibly discordant genes (|log2 FC| > 1, FDR< 0.05 by LIMMA), including cell-type informative features. These discrepancies reflect technology-specific effects and violate the linear mixing assumption underlying most deconvolution methods. We demonstrate that a protocol-matched in-study reference constructed from only six donors consistently outperforms substantially larger external references, with the advantage that increases at finer cell-type resolution. Moreover, paired samples enable the identification and selective removal of discordant genes, which further improves deconvolution accuracy for many algorithms, particularly in high-resolution settings. These findings are robust across methods, references, and evaluation criteria and extend beyond compositional accuracy to improving recovery of disease-associated cell-type differences in clinical applications. ConclusionsOur study provides the first fully paired benchmark of transcriptomic deconvolution on real bulk RNA-seq data of human BAL samples and demonstrates that reference design and data compatibility are as influential as algorithm choice. Beyond benchmarking, we introduce a practical and cost-effective protocol for deconvolution studies: generate single-cell data for a minimal subset of bulk samples (pilot pairing), use these data to construct an in-study reference and identify discordant genes, and apply the resulting insights to the full cohort. This strategy requires limited additional experimental effort yet yields substantial gains in accuracy and stability, offering actionable guidance for future bulk RNA-seq deconvolution studies across tissues and platforms. One-sentence summarySystematic benchmarking reveals pseudo-bulk biases and provides practical fixes for accurate RNA deconvolution.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- omnideconv: a unifying framework for using and benchmarking single-cell-informed deconvolution of bulk RNA-seq data 97%
- Missing cell types in single-cell references impact deconvolution of bulk data but are detectable 96%
- Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics 96%
Similar papers in this journal
- Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples 95%
- A comprehensive comparison on cell type composition inference for spatial transcriptomics data 94%
- CSRefiner: A lightweight framework for fine-tuning cell segmentation models with small datasets 94%
Similar papers in this journal
- Community assessment of methods to deconvolve cellular composition from bulk gene expression 96%
- On the discovery of population-specific state transitions from multi-sample multi-condition single-cell RNA sequencing data 95%
- Atlas-scale single-cell multi-sample multi-condition data integration using scMerge2 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.