Back

Unsupervised single-cell clustering with Asymmetric Within-Sample Transformation and per cluster supervised features selection

Pagnotta, S. M.

2023-05-21 bioinformatics
10.1101/2023.05.17.541148 bioRxiv
Show abstract

This chapter shows applying the Asymmetric Within-Sample Transformation [14] to single-cell RNA-Seq data matched with a previous dropout imputation. The asymmetric transformation is a special winsorization that flattens low-expressed intensities and preserves highly expressed gene levels. Before a standard hierarchical clustering algorithm, an intermediate step removes non-informative genes according to a threshold applied to a per-gene entropy estimate. Following the clustering, a time-intensive algorithm is shown to uncover the molecular features associated with each cluster. This step implements a resampling algorithm to generate a random baseline to measure up/down-regulated significant genes. To this aim, we adopt a GLM model [10] as implemented in DESeq2 [9] package. We render the results in graphical mode. While the tools are standard heat maps, we introduce some data scaling so that the results reliability is crystal clear.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.