Back

The Interplay Between Sketching and Graph Generation Algorithms in Identifying Biologically Cohesive Cell-Populations in Single-Cell Data

Crawford, E. B.; Plotkin, A.; Ranek, J.; Stanley, N.

2023-09-17 bioinformatics
10.1101/2023.09.15.557825 bioRxiv
Show abstract

High-throughput single-cell immune profiling technologies, such as mass cytometry (CyTOF) and single-cell RNA sequencing measure the expression of multiple proteins or genes across many individual cells within a profiled sample. As it is often of interest to identify particular clusters or cell-populations driving clinical phenotypes or experimental outcomes, there is a critical need to develop automated bioinformatics approaches that can handle a large number of profiled cells. For analyzing multi-sample single-cell datasets at scale, the datasets are usually encoded as a graph, where nodes represent cells and edges imply significant between-cell similarity. As multi-sample single-cell experiments can readily result in millions of profiled cells, the construction and analysis of a graph becomes computationally prohibitive and often requires reducing the dataset size through downsampling as a pre-processing step. Here, we explore the interplay between sketching, or downsampling approaches, and the way in which the graph is constructed on the sketched data for ultimately identifying biologically-meaningful cell-populations. Our results suggest that combining a principled sketching approach with a simple k-nearest neighbor graph representation of the data can identify meaningful subsets of cells as robustly as, and sometimes better than, more sophisticated graph generation approaches. This reveals that the practical concern of downsampling or sketching a limited number of cells is a more critical pre-processing step than how the graph representation is constructed.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.