Unsupervised single-cell clustering with Asymmetric Within-Sample Transformation and per cluster supervised features selection
Pagnotta, S. M.
Show abstract
This chapter shows applying the Asymmetric Within-Sample Transformation [14] to single-cell RNA-Seq data matched with a previous dropout imputation. The asymmetric transformation is a special winsorization that flattens low-expressed intensities and preserves highly expressed gene levels. Before a standard hierarchical clustering algorithm, an intermediate step removes non-informative genes according to a threshold applied to a per-gene entropy estimate. Following the clustering, a time-intensive algorithm is shown to uncover the molecular features associated with each cluster. This step implements a resampling algorithm to generate a random baseline to measure up/down-regulated significant genes. To this aim, we adopt a GLM model [10] as implemented in DESeq2 [9] package. We render the results in graphical mode. While the tools are standard heat maps, we introduce some data scaling so that the results reliability is crystal clear.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Stardust: improving spatial transcriptomics data analysis through space aware modularity optimization based clustering. 96%
- AlcoR: alignment-free simulation, mapping, and visualization of low-complexity regions in biological data 95%
- parSMURF, a High Performance Computing tool for the genome-wide detection of pathogenic variants 95%
Similar papers in this journal
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 95%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 95%
- scHiCTools: a computational toolbox for analyzing single cell Hi-C data 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.