From microbial diversity to function; evaluating dimensionality reduction methods
Chamberlain, E. J.; Boulton, W.; Connors, E.; Calianos, T.; Bowman, J.; Creamean, J.; Mock, T.; Kim, H. H.
Show abstract
Artificial Intelligence (AI), and more specifically Machine Learning (ML), have become an increasingly prevalent tool in microbial oceanography. The high dimensionality of microbial diversity data from omics observations is highly suitable for ML analysis, with many recent studies showcasing their utility for exploratory ecological feature finding and process prediction. Here, we apply three well-documented dimensionality reduction methods including Principal Coordinate Analysis (PCoA), Self Organizing Maps (SOM), and Weighted Gene Correlation Network Analysis (WGCNA), to near daily 16S rRNA gene amplicon sequencing data from the 2019-2020 MOSAiC International Arctic Drift Expedition. We compare the k-means clustering outputs from these methods to extract functionally distinct seasonal microbial ecotypes in the surface Arctic Ocean. Our results indicate the SOM method outperforms a more traditional PCoA ordination, identifying a greater number of metabolically distinct functional groups. We then investigate the importance of including biological context in dimensionality reduction by comparing functional outputs to a taxa clustering approach using a k-means adapted WGCNA correlation network. Regardless of data input, all 3 methods identified 3-4 recurrent ecotypes with distinct taxonomic and functional cut-offs driven by seasonality, water mass, and substrate turnover. Ultimately, these results reinforce such methodologies as a meaningful translator in the mining of historical amplicon datasets to address modern mechanistic questions and incorporate greater ecotype diversity into mechanistic biogeochemical models. ImportanceConnecting microbial community structure to ecosystem function is an important step in accurately modeling climate-relevant biogeochemical processes yet remains a major challenge in microbial oceanography. This manuscript demonstrates how emerging machine learning approaches can establish this connection by uncovering recurrent ecological patterns in Arctic Ocean microbial communities. Using near-daily 16S rRNA gene and supplementary metagenome data from the MOSAiC drift expedition, we identified distinct "ecotypes," or groups of microbes that perform differentiable functional roles within the ecosystem. Importantly, our methods reveal new connections between microbial identity and function that traditional analyses may overlook. It is possible such techniques could be applied to historical amplicon datasets, allowing scientists to revisit and reinterpret existing data to better understand how polar ecosystems are responding to environmental change and to improve future predictive climate models.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Picophytoplankton Implicated in Productivity and Biogeochemistry in the North Pacific Transition Zone 97%
- Quantification of metabolic niche occupancy dynamics in a Baltic Sea bacterial community 97%
- Dominant nitrogen metabolisms of a warm, seasonally anoxic freshwater ecosystem revealed using genome resolved metatranscriptomics 96%
Similar papers in this journal
- Meta-omics reveals role of photosynthesis in Microbially Induced Carbonate Precipitation at a CO2-rich Geyser 96%
- Syndiniales parasites drive species networks and are a biomarker for carbon export in the oligotrophic ocean 96%
- Sea-ice melt determines seasonal phytoplankton dynamics and delimits the habitat of temperate Atlantic taxa as the Arctic Ocean atlantifies 95%
Similar papers in this journal
- The dynamic trophic architecture of open-ocean protist communities revealed through machine- guided metatranscriptomics 96%
- Species invasions shift microbial phenology in a two-decade freshwater time series 96%
- Diatom Modulation of Microbial Consortia Through Use of Two Unique Secondary Metabolites 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.