Whole Metagenome Sequencing: not Deep Enough for Complete Microbial Function Recovery
Liu, J.; Coker, M. O.; Osazuwa-Peters, N.; Peter, O.; Idemudia, N. L.; Schlecht, N. F.; Obuekwe, O.; Eki-Udoko, F. E.; Bromberg, Y.
Show abstract
BackgroundWhole metagenome shotgun sequencing (WMS) is widely used to profile microbial function. However, technical variability in sequencing and analysis often obscures true biological patterns. Large-scale studies are particularly susceptible to batch effects, such as differences in sequencing depth and platform and annotation strategies, as well as sample-to-flow-cell assignments. However, the relative effects of these factors on functional inference in such studies have yet to be systematically evaluated. We analyzed oral-rinse WMS data from a study cohort including 671 Nigerian youths aged 9-18, sequenced on two Illumina platforms. Microbial molecular functionality encoded in these data were annotated using the mi-faser/Fusion pipeline, to capture the broad functional repertoire, and HUMAnN 3/EC numbers pipeline to characterize curated enzymatic activities. We then quantified how technical factors and batch effects shaped the recovery of microbial functionality. ResultsThree findings of our work were most salient. First, we observed that the choice of annotation strategy traded off between breadth and specificity of functional coverage. Second, we found that low-prevalence functions were disproportionately lost at shallow sequencing depths, indicating that in e.g. case-control studies with few representatives of the minor class, sequencing depth could critically impact study resolution. Finally, using our newly developed model relating sequencing depth to functional recovery, we demonstrated that increasing sequencing depth does not directly or proportionally improve functional recall. That is, at as little as 10% of this studys sequencing depth, 30% of the estimated complete microbiome functional repertoire was detectable. However, even at the full depth used in this study, we were only able to recover an estimated 60% of that complete functional repertoire. ConclusionsTogether, these findings and our depth-to-function mapping framework provide practical guidelines for the design and interpretation of WMS studies. Coordinating sequencing depth planning with annotation strategy, experimental design, and rigorous batch control is thus essential for robust detection of microbial functions and for ensuring reproducible microbiome insights.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MetaPro: A scalable and reproducible data processing and analysis pipeline for metatranscriptomic investigation of microbial communities 96%
- Single Amplified Genome Catalog Reveals the Dynamics of Mobilome and Resistome in the Human Microbiome 95%
- Improved eukaryotic detection compatible with large-scale automated analysis of metagenomes 95%
Similar papers in this journal
- Illumina Complete Long Read Assay yields contiguous bacterial genomes from human gut metagenomes 96%
- BiG-MAP: an automated pipeline to profile metabolic gene cluster abundance and expression in microbiomes 95%
- Longitudinal, Multi-platform Metagenomics Yields a High-quality Genomic Catalog and Guides an In Vitro Model for Cheese Communities 95%
Similar papers in this journal
- Evaluating de novo assembly and binning strategies for time-series drinking water metagenomes. 94%
- Utilizing co-abundances of antimicrobial resistance genes to identify potential co-selection in the resistome 94%
- Heterogeneous lineage-specific arginine deiminase expression within dental microbiome species 94%
Similar papers in this journal
- Growth phase estimation for abundant bacterial populations sampled longitudinally from human stool metagenomes. 95%
- Taxometer: Improving taxonomic classification of metagenomics contigs 95%
- Systematic evaluation of metatranscriptomic differential gene expression in silico, in vitro, and in vivo enables elucidation of inter-species cross-feeding 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.