Deduplication Improves Cost-Efficiency and Yields of De novo Assembly and Binning of Shot-Gun Metagenomes in Microbiome Research
Zhang, Z.; Zhang, L.; Zhao, Z.; Wang, H.; Ju, F.
Show abstract
Metagenomics has in the last decade greatly revolutionized the study of microbial communities. However, the presence of artificial duplicate reads mainly raised from the preparation of metagenomic DNA sequencing library and their impacts on metagenomic assembly and binning have never brought to the attention. Here, we explicitly investigated the effects of duplicate reads on metagenomic assembly and binning, based on analyses of four groups of representative metagenomes with distinct microbiome complexity. Our results showed that deduplication considerably increased the binning yields (by 3.5% to 80%) for most of the metagenomic datasets examined thanks to improved contig length and coverage profiling of metagenome-assembled contigs. Specifically, 411 versus 397, 331 versus 317, 104 versus 88 and 9 versus 5 metagenome-assembled genomes (MAGs) were recovered from MEGAHIT assemblies of bioreactor sludge, surface water, lake sediment, and forest soil metagenomes, respectively. Noticeably, deduplication reduced the computational costs of metagenomic assembly including elapsed time (by 9.0% to 29.9%) and maximum memory requirement (by 4.3% to 37.1%). Collectively, it is recommended to remove duplicate reads in metagenomic data before assembly and binning analyses, particularly for complex environmental samples, such as forest soils examined in this study. ImportanceDuplicated reads are usually considered as technical artefacts. Their presence in metagenomes would theoretically not only introduce bias in the quantitative analysis, but also result in mistakes in coverage profile, leading to negative effects or even failures on metagenomic assembly and binning, as the widely used metagenome assemblers and binners all need coverage information for graph partitioning and assembly binning, respectively. However, this issue was seldomly noticed and its impacts on the downstream key bioinformatic procedures (e.g., assembly and binning) still remained unclear. In this study, we comprehensively evaluated for the first time the impacts of duplicate reads on de novo assembly and binning of real metagenomic datasets by comparing assembly quality, binning yields and the requirements of computational resources with and without the removal of duplicate reads. It was revealed that deduplication considerably increased the binning yields and significantly reduced the computational costs including elapsed time and maximum memory requirement. The results provide empirical reference for more cost-efficient metagenomic analyses in microbiome research.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Phages-bacteria interactions underlying the dynamics of polyhydroxyalkanoates-producing mixed microbial cultures via meta-omics study 95%
- Genome and community-level interaction insights on wide carbon utilizing and element cycling function of Hydrothermarchaeota from hydrothermal sediment 95%
- Comparative genomic insights into the evolution of Halobacteria-associated "Candidatus Nanohaloarchaeota" 95%
Similar papers in this journal
- Evaluating de novo assembly and binning strategies for time-series drinking water metagenomes. 95%
- Library Preparation and Sequencing Platform Introduce Bias in Metagenomic-Based Characterizations of Microbiomes 94%
- Living to the high extreme: unraveling the composition, structure, and functional insights of bacterial communities thriving in the arsenic-rich Salar de Huasco - Altiplanic ecosystem. 94%
Similar papers in this journal
Similar papers in this journal
- Exploring Taxonomic and Functional Microbiome of Hawaiian Stream and Spring Irrigation Water Systems Using Illumina and Oxford Nanopore Sequencing Platforms 96%
- Metagenomic analysis of ecological niche overlap and community collapse in microbiome dynamics 95%
- Metagenomic association analysis of gut symbiont Lactobacillus reuteri without host-specific genome isolation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.