Back

Refined Enterotyping Reveals Dysbiosis in Global Fecal Metagenomes

Keller, M. I.; Nishijima, S.; Podlesny, D.; Kim, C. Y.; Robbani, S. M.; Schudoma, C.; Fullam, A.; Richter, J.; Letunic, I.; Akanni, W.; Orakov, A.; Schmidt, T. S.; Marotta, F.; Trebicka, J.; Kuhn, M.; Van Rossum, T.; Bork, P.

2024-08-13 microbiology
10.1101/2024.08.13.607711 bioRxiv
Show abstract

BackgroundEnterotypes describe human fecal microbiomes grouped by similarity into clusters of microbial community composition, often associated with disease, medications, diet, and lifestyle. Numbers and determinants of enterotypes have been derived by diverse frameworks and applied to cohorts that often lack diversity or inter-cohort comparability. ResultsTo overcome these limitations, we selected 16,772 fecal metagenomes collected from 38 countries to revisit the enterotypes using state-of-the-art fuzzy clustering and found robust clustering regardless of underlying taxonomy, consistent with previous findings. Quantifying the strength of enterotype classifications enriched the enterotype landscape, also reflecting some continuity of microbial compositions. As the classification strength was associated with the patients health status, we established an "Enterotype Dysbiosis Score" (EDS) as a latent covariate for various diseases. ConclusionThis global study confirms the enterotypes, reveals a dysbiosis signal within the enterotype landscape, and enables robust classification of metagenomes with an online "Enterotyper" tool, allowing reproducible analysis in future studies. Graphical Abstract O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/607711v3_ufig1.gif" ALT="Figure 1"> View larger version (30K): org.highwire.dtl.DTLVardef@1611a2org.highwire.dtl.DTLVardef@dfbb57org.highwire.dtl.DTLVardef@848ac0org.highwire.dtl.DTLVardef@1b15808_HPS_FORMAT_FIGEXP M_FIG C_FIG

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.