Back

Integrating single-cell and single-nucleus datasets improves bulk RNA-seq deconvolution

Ivich, A.; Greene, C. S.

2025-08-24 bioinformatics
10.1101/2025.08.20.671333 bioRxiv
Show abstract

Bulk RNA-seq deconvolution typically uses single-cell RNA-sequencing (scRNA-seq) references, but some cell types are only detectable through single-nucleus RNA sequencing (snRNA-seq). Because snRNA-seq captures nuclear, but not cytoplasmic, transcripts, direct use as a reference could reduce deconvolution accuracy. Here, we systematically benchmark strategies to integrate both modalities, focusing on transformations and gene-filtering approaches that harmonize snRNA-seq with scRNA-seq references. Across four diverse tissues, we evaluated principal component-based shifts, conditional and non-conditional variational autoencoders (scVI), and the removal of cross-modality differentially expressed genes (DEGs). While all methods improved performance relative to untransformed snRNA-seq, filtering consistent cross-modality DEGs delivered the greatest gains, often matching or surpassing scRNA-only references. Conditional scVI performed comparably and was especially effective when matched scRNA-snRNA cell types were unavailable. In real adipose bulk samples without ground truth, DEG pruning and conditional scVI provided the most robust cell-fraction estimates across donors and transformations. Together, these results demonstrate that scRNA-seq should be prioritized as the reference when available, with snRNA-seq appended only after filtering cross-modality DEGs. For less-characterized systems where DEG information is limited, conditional scVI offers a practical alternative. Our findings provide clear guidelines for modality-aware integration, enabling near-scRNA-seq accuracy in bulk deconvolution workflows.

Published in Cell Reports Methods (predicted rank #8) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.