Back

Mild CF Lung Disease is Associated with Bacterial Community Stability

Hampton, T. H.; Thomas, D.; van der Gast, C.; Stanton, B. A.; O'Toole, G.

2021-03-24 microbiology
10.1101/2021.03.23.436717 bioRxiv
Show abstract

Microbial communities in the airways of persons with CF (pwCF) are variable, may include genera that are not typically associated with CF, and their composition can be difficult to correlate with long-term disease outcomes. Leveraging two large datasets characterizing sputum communities of 167 pwCF and associated metadata, we identify five bacterial community types. These communities explain 24% of the variability in lung function in this cohort, far more than single factors like Simpson diversity, which explains only 4%. Subjects with Pseudomonas-dominated communities tended to be older and have reduced percent predicted FEV1 (ppFEV1) than subjects with Streptococcus-dominated communities, consistent with previous findings. To assess the predictive power of these five communities in a longitudinal setting, we used random forests to classify 346 additional samples from 24 subjects observed 8 years on average in a range of clinical states. Subjects with mild disease were more likely to be observed at baseline, that is, not in the context of a pulmonary exacerbation, and community structure in these subjects was more self-similar over time, as measured by Bray-Curtis distance. Interestingly, we found that subjects with mild disease were more likely to remain in a mixed Pseudomonas community, providing some support for the climax-attack model of the CF airway. In contrast, patients with worse outcomes were more likely to show shifts among community types. Our results suggest that bacterial community instability may be a risk factor for lung function decline and indicates the need to better understand factors that drive shifts in community composition.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.