Back

k-spaces: Mixtures of Gaussian latent variable models

Markarian, N.; Engelhardt, B. E.; Pierce, N. A.; Sternberg, P. W.; Pachter, L.

2025-11-28 bioinformatics
10.1101/2025.11.24.690254 bioRxiv
Show abstract

Principal component analysis (PCA) and k-means clustering are two seemingly different methods for dimension reduction and clustering, respectively, but can be understood as special cases of inference in a Gaussian latent variable model framework. We leverage this insight to develop a probabilistic framework and methods for simultaneous dimension reduction, clustering, and latent space learning that are efficient and interpretable, and that can replace current ad hoc combinations of PCA and clustering. The algorithm, k-spaces, has broad applicability, which we demonstrate in several distinct genomic settings. In particular, we show how k-spaces can be used to model gene expression in quantitative hybridization chain reaction (qHCR) images, for inference in epigenomics, and for dimension reduction of single-cell RNA-sequencing data.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.