Back

Fast and interpretable scRNA-seq data analysis

Cobanoglu, M. C.

2020-10-07 bioinformatics
10.1101/2020.10.05.314039 bioRxiv
Show abstract

One of the key challenges in single-cell data analysis is the annotation of cells with their cell types. This task is divided into two different sub-tasks: identifying known cell types and identifying novel cell types. In the former case, we can benefit from being able to transfer annotations from bulk RNA-seq because there are many more types profiled with that more established technology. In the latter case, we would benefit from interpretable models that can describe the reasons for grouping a number of cells together. We propose that both of these problems can be solved by generative Bayesian Dirichlet-multinomial models. In the supervised learning context, we propose a generative Bayesian Dirichlet-multinomial classifier. We show that such a classifier can effectively transfer cell labels from bulk to single-cell RNA-sequencing data. We also show that alternative well-established machine learning models have difficulty with this transition, even if they are effective within the same regime (i.e. single cell to single cell). In the unsupervised learning context, we propose a Bayesian Dirichlet-multinomial mixture model. We show that the proposed model learns meaningful clusters where the automatically learned relationships between cell types and genes overlap with ground truth associations. Furthermore, there are no density or connectivity based clustering assumptions in this model, which differs with almost every approach in this field. Consequently the clustering results from the generative method can effectively represent nuanced differences among cells. Contactmurat.cobanoglu@utsouthwestern.edu

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.