Back

A hyperspherical deep Bayesian model for interpretable clustering and relationship prediction in microbiome multi-omics integration

Dang, T.; Lysenko, A.; Tsunoda, T.

2026-08-01 bioinformatics
10.64898/2026.07.28.741374 bioRxiv
Show abstract

The microbiome plays a significant role in the development and progression of many diseases, yet extracting interpretable insights from multi-omics data remains challenging. Existing approaches face a recurring practical trade-off: deep learning methods achieve high predictive performance but lack uncertainty quantification, whereas probabilistic methods provide interpretable results but require data-type-specific likelihood functions that limit generalization across diverse omics modalities. Here, we introduce DBayesCM (Deep Bayesian Clustering for Multi-omics), which combines deep learning modeling with Bayesian nonparametric methods. DBayesCM employs separate encoders to project microbiome and host omics data into a shared latent space, where an infinite mixture model with a Dirichlet process prior determines the number of clusters automatically while quantifying the uncertainty of each samples assignment. Spike-and-slab priors identify discriminative features, and a Bayesian neural network estimates probabilistic co-occurrence between microbial species and host omics features. To isolate the effect of latent geometry, we evaluate two variants that are identical except for their latent space: DBayesCM-vMF constrains the latent to the unit hypersphere and applies a von Mises-Fisher mixture, while DBayesCM-GMM uses a Euclidean latent space and a Gaussian mixture. On simulated data, the hyperspherical variant recovered the correct number of clusters, whereas the Euclidean variant over-segmented, demonstrating that the latent geometry affects cluster recovery. Applied to colon, breast, and kidney cancer cohorts spanning metagenomics, host metabolomics, RNA-seq, and miRNA data, and to an obstructive sleep apnea model, DBayesCM ranked consistently among the existing methods while uniquely combining data-driven cluster-number determination, sample-level uncertainty, and interpretable feature selection within a single framework. DBayesCM reveals conditional probabilistic co-occurrence between core microbial species and host omics features, enabling uncertainty-aware exploration of microbiome-host relationships across diverse diseases.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.