Back

HIDE: Hierarchical cell-type Deconvolution

Völkl, D.; Mensching-Buhr, M.; Sterr, T.; Bolz, S.; Schäfer, A.; Seifert, N.; Tauschke, J.; Rayford, A.; Straume, O.; Zacharias, H. U.; Grellscheid, S. N.; Beissbarth, T.; Altenbuchinger, M.; Görtler, F.

2025-02-05 bioinformatics
10.1101/2025.01.31.634483 bioRxiv
Show abstract

MotivationCell-type deconvolution is a computational approach to infer cellular distributions from bulk transcriptomics data. Several methods have been proposed, each with its own advantages and disadvantages. Reference based approaches make use of archetypic transcriptomic profiles representing individual cell types. Those reference profiles are ideally chosen such that the observed bulks can be reconstructed as a linear combination thereof. This strategy, however, ignores the fact that cellular populations arise through the process of cellular differentiation, which entails the gradual emergence of cell groups with diverse morphological and functional characteristics. ResultsHere, we propose Hierarchical cell-type Deconvolution (HIDE), a cell-type deconvolution approach which incorporates a cell hierarchy for improved performance and interpretability. This is achieved by a hierarchical procedure that preserves estimates of major cell populations while inferring their respective subpopulations. We show in simulation studies that this procedure produces more reliable and more consistent results than other state-of-the-art approaches. Finally, we provide an example application of HIDE to explore breast cancer specimens from TCGA. AvailabilityA python implementation of HIDE is available at zenodo (doi:10.5281/zenodo.14724906). Supplementary informationSupplementary material is available at Bioinformatics online.

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.