Back

FastPG: Fast clustering of millions of single cells

Bodenheimer, T.; Halappanavar, M.; Jefferys, S.; Gibson, R.; Liu, S.; Mucha, P. J.; Stanley, N.; Parker, J. S.; Selitsky, S. R.

2020-06-20 bioinformatics
10.1101/2020.06.19.159749 bioRxiv
Show abstract

Current single-cell experiments can produce datasets with millions of cells. Unsupervised clustering can be used to identify cell populations in single-cell analysis but often leads to interminable computation time at this scale. This problem has previously been mitigated by subsampling cells, which greatly reduces accuracy. We built on the graph-based algorithm PhenoGraph and developed FastPG which has the same cell assignment accuracy but is on average 27x faster in our tests. FastPG also has higher cell assignment accuracy than two other fast clustering methods, FlowSOM and PARC. AvailabilityFastPG is available here: https://github.com/sararselitsky/FastPG

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.