Back

Phenotype-driven parallel embedding for microbiome multi-omic data integration

Bamberger, T.; consortium, D.; Borenstein, E.

2025-12-06 bioinformatics
10.64898/2025.12.05.692119 bioRxiv
Show abstract

The human microbiome is widely recognized as a key determinant of health and disease, yet most reported links between observed microbial features and clinical outcomes remain descriptive and lack an integrated system-level perspective. Multi-omic studies of the microbiome, which jointly profile and analyze multiple molecular aspects of the microbiome via metagenomics, metabolomics, proteomics, and transcriptomics assays, offers a more comprehensive view of this system, with the potential to uncover how microbial communities and functions influence host physiology. However, integration of such multi-omic data remains challenging due to high dimensionality, major differences in data properties across omics, and the need to utilize and preserve omic-specific information. Embedding omic data in low-dimensional spaces offer a promising avenue to capture complex patterns, reduce noise, and improve downstream analysis, yet most embedding-based microbiome studies to date exhibited limited predictive power or relied on a single joint embedding of all omics thus failing to preserve omic-species properties. To address this, we introduce PAPRICA (Phenotype-Aware Parallel Representation for Integrative omiC Analysis), an encoder-decoder framework for microbiome multi-omic integration that embeds each omic into its own latent space while jointly modeling their relationships. The model consists of parallel autoencoders trained with a loss function that promotes three objectives: (1) accurate reconstruction of each omic, (2) alignment of samples across omics such that proximity in one latent space reflects proximity in the others, and (3) alignment with a phenotype space to capture variation associated with continuous outcomes, such as fecal calprotectin levels in IBD. The resulting models support cross-omic inference and phenotype prediction from the learned latent representations, and enables integration without collapsing data into a single space. This modeling approach thus preserves omic-specific signals while capturing phenotype-associated variation. We compared PAPRICA to four alternative models that represent successive advances in multi-omic integration architectures. We found that across two complementary tasks, predicting one omic profile from another and predicting a continuous phenotype from an input omic profile, our parallel autoencoder approach, and particularly the PAPRICA model, demonstrated better performance across three multi-omic datasets (the Franzosa IBD cohort, Lifelines DEEP and the Dog Aging Project Precision Cohort). Combined, these findings suggest that our embedding strategy effectively captures and balances omic-specific structure, cross-omic relationships, and phenotype-relevant signals across diverse datasets, offering a flexible, scalable framework for embedding complex multi-omic microbiome data and advancing our ability to gain new insights into host-microbiome interactions.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.