Comprehensive benchmarking of computational deconvolution of transcriptomics data
Avila Cobos, F.; Alquicira-Hernandez, J.; Powell, J.; Mestdagh, P.; De Preter, K.
Show abstract
Many computational methods to infer cell type proportions from bulk transcriptomics data have been developed. Attempts comparing these methods revealed that the choice of reference marker signatures is far more important than the method itself. However, a thorough evaluation of the combined impact of data transformation, pre-processing, marker selection, cell type composition and choice of methodology on the results is still lacking. Using different single-cell RNA-sequencing (scRNA-seq) datasets, we generated hundreds of pseudo-bulk mixtures to evaluate the combined impact of these factors on the deconvolution results. Along with methods to perform deconvolution of bulk RNA-seq data we also included five methods specifically designed to infer the cell type composition of bulk data using scRNA-seq data as reference. Both bulk and single-cell deconvolution methods perform best when applied to data in linear scale and the choice of normalization can have a dramatic impact on the performance of some, but not all methods. Overall, single-cell methods have comparable performance to the best performing bulk methods and bulk methods based on semi-supervised approaches showed higher error and lower correlation values between the computed and the expected proportions. Moreover, failure to include cell types in the reference that are present in a mixture always led to substantially worse results, regardless of any of the previous choices. Taken together, we provide a thorough evaluation of the combined impact of the different factors affecting the computational deconvolution task across different datasets and propose general guidelines to maximize its performance.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Systematic evaluation of transcriptomics-based deconvolution methods and references using thousands of clinical samples 96%
- IBRAP: Integrated Benchmarking Single-cell RNA-sequencing Analytical Pipeline 95%
- SCDC: Bulk Gene Expression Deconvolution by Multiple Single-Cell RNA Sequencing References 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Integrating single-cell and single-nucleus datasets improves bulk RNA-seq deconvolution 94%
- UniFORM: Towards Universal Immunofluorescence Normalization for Multiplex Tissue Imaging 94%
- SELINA: Single-cell Assignment using Multiple-Adversarial Domain Adaptation Network with Large-scale References 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.