Back

Subspace clustering identifies transcriptome constraints determining cell type identity

Huang, A.; Kim, J.

2025-12-04 cell biology
10.64898/2025.12.04.692451 bioRxiv
Show abstract

With the advent of single-cell RNA-sequencing, researchers now have the ability to define cell types from large amounts of transcriptome information. Currently, most clustering algorithms measure cell-to-cell similarities using distance metrics based on the assumption that each cluster is comprised of "nearby" neighbors. In effect, clusters are a collection of similar cells in the embedded metric. Here, we propose that biological clusters should be comprised of sets of cells that satisfy a set of stochiometric constraints, whose intersections define a cell type. We propose to model each cell population with a single affine subspace, where all cells of the same type share a common set of linear constraints. We present an algorithm that leverages this subspace structure and learns a cell-to-cell affinity matrix based on notions of subspace similarity. We simulate scRNA-seq data according to the subspace model and benchmark our algorithm against pre-existing methods. We further benchmark our algorithm on a C. elegans dataset and show recovery of information on both cell type and developmental time. Lastly, we find the subspaces that our algorithm recovers allow us to find biologically significant genes involved in an organisms development.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.