Phenotype-driven parallel embedding for microbiome multi-omic data integration
Bamberger, T.; consortium, D.; Borenstein, E.
Show abstract
The human microbiome is widely recognized as a key determinant of health and disease, yet most reported links between observed microbial features and clinical outcomes remain descriptive and lack an integrated system-level perspective. Multi-omic studies of the microbiome, which jointly profile and analyze multiple molecular aspects of the microbiome via metagenomics, metabolomics, proteomics, and transcriptomics assays, offers a more comprehensive view of this system, with the potential to uncover how microbial communities and functions influence host physiology. However, integration of such multi-omic data remains challenging due to high dimensionality, major differences in data properties across omics, and the need to utilize and preserve omic-specific information. Embedding omic data in low-dimensional spaces offer a promising avenue to capture complex patterns, reduce noise, and improve downstream analysis, yet most embedding-based microbiome studies to date exhibited limited predictive power or relied on a single joint embedding of all omics thus failing to preserve omic-species properties. To address this, we introduce PAPRICA (Phenotype-Aware Parallel Representation for Integrative omiC Analysis), an encoder-decoder framework for microbiome multi-omic integration that embeds each omic into its own latent space while jointly modeling their relationships. The model consists of parallel autoencoders trained with a loss function that promotes three objectives: (1) accurate reconstruction of each omic, (2) alignment of samples across omics such that proximity in one latent space reflects proximity in the others, and (3) alignment with a phenotype space to capture variation associated with continuous outcomes, such as fecal calprotectin levels in IBD. The resulting models support cross-omic inference and phenotype prediction from the learned latent representations, and enables integration without collapsing data into a single space. This modeling approach thus preserves omic-specific signals while capturing phenotype-associated variation. We compared PAPRICA to four alternative models that represent successive advances in multi-omic integration architectures. We found that across two complementary tasks, predicting one omic profile from another and predicting a continuous phenotype from an input omic profile, our parallel autoencoder approach, and particularly the PAPRICA model, demonstrated better performance across three multi-omic datasets (the Franzosa IBD cohort, Lifelines DEEP and the Dog Aging Project Precision Cohort). Combined, these findings suggest that our embedding strategy effectively captures and balances omic-specific structure, cross-omic relationships, and phenotype-relevant signals across diverse datasets, offering a flexible, scalable framework for embedding complex multi-omic microbiome data and advancing our ability to gain new insights into host-microbiome interactions.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- LEOPARD: missing view completion for multi-timepoints omics data via representation disentanglement and temporal knowledge transfer 97%
- Multi-omic integration of microbiome data for identifying disease-associated modules 97%
- Ecology-guided prediction of cross-feeding interactions in the human gut microbiome 96%
Similar papers in this journal
- APOLLO: A genome-scale metabolic reconstruction resource of 247,092 diverse human microbes spanning multiple continents, age groups, and body sites 95%
- Integration of multi-modal measurements identifies critical mechanisms of tuberculosis drug action 93%
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 93%
Similar papers in this journal
- Benchmarking Differential Abundance Analysis Methods for Correlated Microbiome Sequencing Data 95%
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 94%
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.