Back

Gradient boosting regression and convolution improve deconvolution of bulk transcriptomes

Wolfram-Schauerte, M.; Vogel, T.; Achauer, L.; Faelth Savitski, S. M.; Tuoken, H.; Simon, E.; Nieselt, K.

2026-03-05 bioinformatics
10.64898/2026.03.03.709368 bioRxiv
Show abstract

Bulk cell type deconvolution aims to estimate cell type composition from bulk transcriptomic data. So-called pseudobulk simulation, where single-cell RNA-seq data is aggregated to bulk-like expression profiles, represents a central concept in training and testing of deconvolution tools. However, deconvolution methods often lack interpretability and struggle to generalize from simulated to real data. We present GrooD (GradientBoostedDeconvolution), a second-generation deconvolution tool that uses gradient boosted trees trained on pseudobulks from scRNA-seq references. GrooDs pseudobulk simulations account for donor and condition variability to better model the complexity of real-world transcriptomes. We show that GrooD achieves state-of-the-art and superior deconvolution performance on human blood biospecimen. Furthermore, Grood provides visualizations of feature loadings and deconvolution results, that allow mechanistic insights for biological interpretation. We further integrate a convolution framework to assess the transcriptomic similarity between bulk and pseudobulk data, showing that higher similarity can indicate better deconvolution performance. GrooD deconvolution of blood transcriptomes from a large sepsis patient cohort identifies meaningful shifts in immune cell type composition that are associated with disease severity. By combining interpretability, robustness, and heterogeneous pseudobulk simulation, GrooD represents a powerful, user-friendly second-generation tool for cell type deconvolution.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.