Back

Saccharomycotina yeasts defy longstanding macroecological patterns

David, K. T.; Harrison, M.-C.; Opulente, D. A.; LaBella, A. L.; Wolters, J. F.; Zhou, X.; Shen, X.-X.; Groenewald, M.; Pennell, M.; Hittinger, C. T.; Rokas, A.

2023-08-31 ecology
10.1101/2023.08.29.555417 bioRxiv
Show abstract

The Saccharomycotina yeasts ("yeasts" hereafter) are a fungal clade of scientific, economic, and medical significance. Yeasts are highly ecologically diverse, found across a broad range of environments in every biome and continent on earth1; however, little is known about what rules govern the macroecology of yeast species and their range limits in the wild2. Here, we trained machine learning models on 12,221 occurrence records and 96 environmental variables to infer global distribution maps for 186 yeast species ([~]15% of described species from 75% of orders) and to test environmental drivers of yeast biogeography and macroecology. We found that predicted yeast diversity hotspots occur in mixed montane forests in temperate climates. Diversity in vegetation type and topography were some of the greatest predictors of yeast species richness, suggesting that microhabitats and environmental clines are key to yeast diversification. We further found that range limits in yeasts are significantly influenced by carbon niche breadth and range overlap with other yeast species, with carbon specialists and species in high diversity environments exhibiting reduced geographic ranges. Finally, yeasts contravene many longstanding macroecological principles, including the latitudinal diversity gradient, temperature-dependent species richness, and latitude-dependent range size (Rapoports rule). These results unveil how the environment governs the global diversity and distribution of species in the yeast subphylum. These high-resolution models of yeast species distributions will facilitate the prediction of economically relevant and emerging pathogenic species under current and future climate scenarios.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.