Back

Clustering of Omic Data Using Semi-Supervised Transfer Learning for Gaussian Mixture Models via Natural-Gradient Variational Inference: Method and Applications to Bulk and Single-Cell Transcriptomics

Jia, Q.; Conti, D. V.; Goodrich, J. A.

2025-11-14 bioinformatics
10.1101/2025.11.13.688299 bioRxiv
Show abstract

Recent advances in high-throughput technologies have enabled observational studies to collect high-dimensional omic data. However, such data, often measured on small sample sizes, pose challenges to model-based clustering approaches such as Gaussian Mixture Models. Existing methods often fail to generalize due to model instability under complex mixture patterns. To overcome these limitations, we propose a natural-gradient variational inference framework for Gaussian mixture models named Praxis-BGM that incorporates informative priors--cluster-specific means, covariances, and structural connectivity--from large-scale reference data with known cluster or class labels to enable semi-supervised transfer learning. We derive natural-gradient updates that integrate prior knowledge, leveraging the Variational Online Newton algorithm. We also perform feature selection for clustering using Bayes Factors. Implemented using the JAX library for accelerator-oriented computation, Praxis-BGM is computationally efficient and scalable. We demonstrate the effectiveness of Praxis-BGM in extensive simulations and with two real-world applications: bulk transcriptomic datasets for breast cancer subtyping (the Cancer Genome Atlas Breast Invasive Carcinoma and the Molecular Taxonomy of Breast Cancer International Consortium), and transferring cell-type annotations between single-cell transcriptomic datasets produced by different single-cell RNA-seq technologies in a human pancreas study. Even when priors are partially mismatched with the target data, Praxis-BGM enhances semi-supervised clustering accuracy and biological interpretability.

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.