Profiling ranked list enrichment scoring in sparse data elucidates algorithmic tradeoffs
Wenzel, A. T.; Jun, J.; Liefeld, T.; Tamayo, P.; Mesirov, J. P.
Show abstract
Gene Set Enrichment Analysis (GSEA) is a method for quantifying pathway and process activation in groups of samples, and its single sample version (ssGSEA) scores activation using mRNA abundance in a single sample. GSEA and ssGSEA were developed for "bulk" samples rather than individual cell technologies such as microarrays and bulk RNA-sequencing (RNA-seq) data. The growing use of single cell RNA-sequencing (scRNA-seq) raises the possibility of using ssGSEA to quantify pathway and process activation in individual cells. However, scRNA-seq data is much sparser than RNA-seq data. Here we show that ssGSEA as designed for bulk data is subject to some amount of score uncertainty and other technical issues when applied to individual cells from scRNA-seq data. We also show that a ssGSEA can be applied robustly to "pseudobulk" aggregate groups of a few hundred to a few thousand cells provided appropriate normalization is used. Finally, in comparing this approach to other ranked list enrichment methods, we find that the UCell method is most robust to sparsity. We have made the aggregate cell version of ssGSEA available as a Python package and GenePattern module and will also modularize UCell for use on GenePattern as well.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 97%
- Benchmarking algorithms for joint integration of unpaired and paired single-cell RNA-seq and ATAC-seq data 96%
- Missing cell types in single-cell references impact deconvolution of bulk data but are detectable 96%
Similar papers in this journal
Similar papers in this journal
- Binomial models uncover biological variation during feature selection of droplet-based single-cell RNA sequencing 96%
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 95%
- G2S3: a gene graph-based imputation method for single-cell RNA sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.