LGTM: Gaussian Process Modulated Neural Topic Modeling for Longitudinal Microbiome
Yuan, X.; Arany, A.; Formanek, A.; Moreau, Y.; Lähdesmäki, H.; Vatanen, T.
Show abstract
Longitudinal microbiome data are key to understanding the dynamics of microbial communities and their relationships with the host and environment. However, analysis of such data is challenging due to high dimensionality, compositionality, irregular sampling and temporal dependencies on external covariates. Existing analytical approaches typically address only subsets of these challenges, limiting their ability to yield biologically interpretable insights. We introduce LGTM, a probabilistic modeling framework that combines flexible non-linear longitudinal modeling with interpretable topic-based representations of the microbiome. LGTM simultaneously discovers coherent microbial subcommunities ("topics") and models how their abundances change over time and in relation to host and environmental covariates. Using multiple longitudinal human gut microbiome datasets, we demonstrate that LGTM identifies diverse and stable microbial topics while achieving competitive performance in imputation and forecasting tasks. A key strength of the framework is its interpretability: LGTM discovers biologically coherent microbial topics and directly quantifies associations between covariates and microbial dynamics. LGTM is available at https://github.com/yuanx749/lgtm.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A randomization-based causal inference framework for uncovering environmental exposure effects on human gut microbiota 94%
- ResMiCo: increasing the quality of metagenome-assembled genomes with deep learning 94%
- Linear-regression-based algorithms can succeed at identifying microbial functional groups despite the nonlinearity of ecological function 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.