Back

Microbiomes without boundaries: cystic fibrosis pulmotype classifications are dependent on algorithm choice and database size, and indicate continuous variation.

Zhao, C. Y.; Lowhorn, R.; Song, H.; Eum, J.; Brown, S. P.

2025-05-08 infectious diseases
10.1101/2025.05.07.25326893 medRxiv
Show abstract

A common response to microbiome sample variation is to use clustering algorithms to reduce complex and variable datasets to a smaller number of types (e.g. enterotypes for gut samples, or pulmotypes for lung samples). In light of recent analyses showing distinct clustering solutions to in principle similar datasets, we examine the extent to which clustering solutions are dependent on researcher choices of algorithm and dataset, using cystic fibrosis (CF) sputum microbiome data as a model system. Following a structured literature review, we identified 36 CF microbiome studies with publicly available samples and metadata. From these studies we curated a dataset of 4026 sputum microbiome samples across 1184 people with CF (pwCF), complete with matched individual metadata, using a standardized bio-informatic platform. Applying multiple clustering algorithms (DMM, k-means, PAM) to cross-sectional data we find that the optimal clustering varies with both algorithm choice and database size, with generally weak separation among clusters in any classification. Our longitudinal data analyses highlight substantial persistence of cluster types in time, with transitions most common among clusters that are structurally similar, reflecting an underlying continuous landscape of microbiome variation. While transitions among similar clusters are common (e.g. along gradients of Pseudomonas aeruginosa relative abundance), transitions are generally bi-directional, with no clear pathogen-dominated end point states. Using samples from 482 pwCF with available lung function data, we find that taxon-based models outperform cluster-based statistical models in predicting clinical lung function data. Together our results highlight that clustering methods can impose arbitrary boundaries on an underlying continuum of microbiome variation. ImportanceClassifying microbiome samples into discrete "types" is a widely used strategy for simplifying complex microbial community data and linking community structure to clinical outcomes. Here we evaluate the utility of cluster-based microbiome typing schemes, using cystic fibrosis (CF) sputum samples as a model system. We conduct a comprehensive re-analysis of over 4000 sputum samples from more than 1000 people with CF. We show that pulmotype classification outcomes are highly sensitive to the choice of clustering algorithm and dataset size, and that clustering can impose artificial boundaries on a continuous landscape of microbial variation. Our findings urge caution over the use of discrete microbiome classifications and emphasize the value of taxon-based models in capturing the ecology and clinical relevance of complex microbial communities.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.