Back

Benchmarking cell-type deconvolution in cross-platform transcriptomic data

Singh, A.; Cakmak, P.; Lun, J. H.; Macas, J.; Plate, K. H.; Reiss, Y.; Schupp, J.; Imkeller, K.

2025-11-26 bioinformatics
10.1101/2025.11.24.690141 bioRxiv
Show abstract

BackgroundTranscriptomic data from diverse measurement technologies are widely used to study tissue heterogeneity. Cell-type deconvolution, which resolves mixed transcriptomic signals into cellular components, is a key analytical approach. However, achieving accurate deconvolution across platforms remains challenging due to platform-specific experimental and technological biases. ResultsWe systematically benchmarked deconvolution performance using real-world cross-platform datasets and simulated data modeling distinct technological features. Our analyses provide practical guidelines for experimental design and promote more robust and comparable cross-platform transcriptomic analyses. ConclusionLog-normal regression methods such as SpatialDecon demonstrated the most reliable and consistent performance across both simulated and experimental settings, establishing a robust framework for accurate cross-platform transcriptomic deconvolution. Moreover, our results highlight potential caveats in deconvolution predictions arising from specific experimental conditions, providing guidance for more informed experimental design and interpretation Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=60 SRC="FIGDIR/small/690141v1_ufig1.gif" ALT="Figure 1"> View larger version (21K): org.highwire.dtl.DTLVardef@585cbdorg.highwire.dtl.DTLVardef@13088c7org.highwire.dtl.DTLVardef@163ea75org.highwire.dtl.DTLVardef@b5ccf7_HPS_FORMAT_FIGEXP M_FIG C_FIG

Published in Genome Biology (predicted rank #1) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.