Back

Deconvolution of omics data in Python with Deconomix -- cellular compositions, cell-type specific gene regulation, and background contributions

Altenbuchinger, M. C.; Mensching-Buhr, M.; Sterr, T.; Seifert, N.; Voelkl, D.; Tauschke, J.; Rayford, A.; Zacharias, H. U.; Grellscheid, S. N.; Beissbarth, T.; Goertler, F.

2024-12-03 bioinformatics
10.1101/2024.11.28.625894 bioRxiv
Show abstract

SummaryGene expression profiles of heterogeneous bulk samples contain signals from multiple cell populations. Studying variations in their composition can help to identify cell populations relevant for disease. Moreover, analyses, such as the identification of differentially expressed genes, can be confounded by cellular composition, as differences in gene expression may arise from both variations in cellular composition and gene regulation. Here, we present Deconvolution of omics data (Deconomix) - a comprehensive toolbox for the cell-type deconvolution of bulk transcriptomics data. Deconomix stands apart from competing solutions with rich functionality and highly efficient implementations. It facilitates (A) the inference of cellular compositions from bulk transcriptomics data, (B) the machine learning-based optimization of gene weights to resolve small cell populations and to disentangle phenotypically related cells, (C) the inference of background contributions which otherwise would deteriorate cell-type deconvolution, and (D) population estimates of cell-type specific gene regulation. AvailabilityDeconomix is available at https://gitlab.gwdg.de/MedBioinf/MedicalDataScience/Deconomix under GPLv3 licensing. The Python package can be easily installed via pip. It comes with a comprehensive documentation of all user-relevant functions and example workflows provided as Jupyter notebooks.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.