Back

Confounder-free Predictive Models for Microbiome-based Host Phenotype Prediction

Monshizadeh, M.; Hong, Y.; Ye, Y.

2025-02-02 bioinformatics
10.1101/2025.01.29.635502 bioRxiv
Show abstract

As in many fields, the presence of confounding effects (or biases) presents a significant challenge in micro-biome research, including using microbiome data to predict host phenotypes. If not properly addressed, confounders can lead to spurious associations, biased predictions and misleading interpretations. One notable example is the medication metformin, which is commonly prescribed to treat type 2 diabetes (T2D) and is known to influence the gut microbiome. In this study, we propose confounder-free predictive models for human phenotype prediction using microbiome data. These models utilize an end-to-end approach within an adversarial min-max optimization framework to derive features that are invariant to confounding factors, while accounting for the intrinsic correlations between confounders and prediction outcomes. We implemented two versions of confounder-free predictors using different network architectures: one based on a fully connected network (referred as FNN CF) and another incorporating prior biological knowledge (referred as MicroKPNN CF). We evaluated our models on microbiome datasets associated with T2D, where metformin acts as a confounder. Our results demonstrate that confounder-free predictors achieve higher accuracy compared to models that do not account for confounders and more effectively identify microbial markers associated with the phenotype, rather than markers influenced by metformin. Between the two confounder-free models, although the prior-knowledge-guided approach showed slightly lower prediction accuracy compared to the fully connected model, it offered greater interpretability, providing additional insights into the underlying biological mechanisms.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.