bulk2scDiff: A Pseudobulk-Conditioned Diffusion Model for Bulk-to-Single-Cell RNASeq Generation
Xiao, J.; Raue, A.
Show abstract
Bulk RNA sequencing remains the predominant profiling strategy for large clinical cohorts, but it aggregates transcriptional signals across cell populations, thereby masking the underlying cellular heterogeneity. Inferring this heterogeneity from existing bulk transcriptomic data could extend large cohort-based studies that have already been profiled, but constitutes an underdetermined inverse problem, as one bulk profile can be compatible with multiple underlying cellular populations. Existing computational deconvolution methods address this problem primarily by estimating cell-type proportions or cell-type-averaged expression profiles rather than resolving expression at the level of individual cells. Here, we present bulk2scDiff, a proof-of-concept conditional diffusion framework that reformulates bulk-to-single-cell inference as conditional generation of single-cell expression profiles from pseudobulk transcriptomic input. We evaluated bulk2scDiff on two cancer single-cell RNA sequencing datasets, breast cancer and acute myeloid leukemia, where pseudobulk profiles were derived from the single-cell data and used as conditioning inputs, with the matched single-cell populations providing ground truth for controlled evaluation. Across both cases, bulk2scDiff closely reconstructed populations from training samples and generated biologically coherent single-cell populations for held-out samples, generalizing most consistently to recurrent immune features. A pseudobulk-swap control further confirmed sample-specific conditioning, with each sample corresponding pseudobulk yielding the closest agreement with its observed population in nearly all cases. Overall, our work establishes the feasibility of conditional diffusion for generating single-cell populations from pseudobulk transcriptomic profiles, providing a foundation for future evaluation with clinical bulk RNA sequencing data.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Integrating temporal single-cell gene expression modalities for trajectory inference and disease prediction 94%
- omnideconv: a unifying framework for using and benchmarking single-cell-informed deconvolution of bulk RNA-seq data 94%
- Multi-scale deep tensor factorization learns a latent representation of the human epigenome 94%
Similar papers in this journal
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 93%
- scGPD: single-cell informed gene panel design for targeted spatial transcriptomics 93%
Similar papers in this journal
- Multi-resolution deconvolution of spatial transcriptomics data reveals continuous patterns of inflammation 94%
- Density-Preserving Data Visualization Unveils Dynamic Patterns of Single-Cell Transcriptomic Variability 93%
- Foundation Model Attributions Reveal Shared Inflammatory Program Across Diseases 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.