Absence of enterotypes in the human gut microbiomes reanalyzed with non-linear dimensionality reduction methods
Bulygin, I.; Shatov, V.; Rykachevsky, A.; Rayko, A.; Bernstein, A.; Burnaev, E.; Gelfand, M. S.
Show abstract
Enterotypes of the human gut microbiome have been proposed to be a powerful prognostic tool to evaluate the correlation between lifestyle, nutrition, and disease. However, the number of enterotypes suggested in the literature ranged from two to four. The growth of available metagenome data and the use of exact, non-linear methods of data analysis challenges the very concept of clusters in the multidimensional space of bacterial microbiomes. Using several published human gut microbiome datasets, we demonstrate the presence of a lower-dimensional structure in the microbiome space, with high-dimensional data concentrated near a low-dimensional non-linear submanifold, but the absence of distinct and stable clusters that could represent enterotypes. This observation is robust with regard to diverse combinations of dimensionality reduction techniques and clustering algorithms.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Over-optimism in unsupervised microbiome analysis: Insights from network learning and clustering 96%
- Learning, Visualizing and Exploring 16S rRNA Structure Using an Attention-based Deep Neural Network 94%
- The stochastic logistic model with correlated carrying capacities reproduces beta-diversity metrics of microbial 94%
Similar papers in this journal
- ADM: Adaptive Graph Diffusion for Meta-Dimension Reduction 95%
- VBayesMM: Variational Bayesian neural network to prioritize important relationships of high-dimensional microbiome multiomics data 95%
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.