EggVio: a user friendly and versatile pipeline for assembly and functional annotation of shallow depth sequenced samples
Bergk Pinto, B. M.; Vogel, T. M.; Larose, C.
Show abstract
1We introduce a homemade pipeline allowing to improve the quality of the metagenomic annotations carried out when using shallow depth metagenomic datasets. The main motivation being to be able to quantify more precisely, with greater certainty, the genes involved in bacterial interactions. The limitation in our experimental design is that we use a sequencing technique with a low throughput (miSeq) compared to the metagenomic standard (hiSeq) because we carry out a fairly large sampling (almost a hundred samples) in time series. This methodological constraint from our study means that the assembly of the sequences is not very exhaustive (less than 50% of the sequences manage to be assembled). In this chapter, we will therefore present a new pipeline designed to specifically deal with such kind of data. We used co-assembly and a sequence annotation strategy in order to recover the sequences that could not be mapped on the assembled contigs. In addition, in order to avoid adding too much noise, when rescuing reads, we have built an algorithm to define a threshold of e-value based on the noise of the sequence annotation learned from sequences mapped in the assembly. We have selected several recent tools known to be effective for assembling, mapping and annotating these data. In addition, this pipeline was also built in order to be very user-friendly in terms of installation. In this idea of reproducibility, accessibility and transparency, we have designed an installation script to allow each user to install each tool required for the pipeline in a simple and reproducible way. Regarding the performances of this pipeline, we were able to show that the expected error rate (False discovery rate) for the annotation was close to 5%. Finally, we also used an actual dataset from a bioremediation site and showed that the representability of the samples seemed much better when we used our pipeline than when we used a classic metagenome assembly strategy.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- L-norepinephrine Induces Community Shift, Oxidative Stress Response, Metabolic Reprogramming, and Virulence Potential in Wastewater Microbiomes 93%
- Microeukaryotic predators shape the wastewater microbiome. 93%
- Long solids retention times and attached growth phase favor prevalence of comammox bacteria in nitrogen removal systems. 93%
Similar papers in this journal
Similar papers in this journal
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 94%
- The DOE JGI Metagenome Workflow 94%
- Decomposing a San Francisco Estuary microbiome using long read metagenomics reveals species and species- and strain-level dominance from picoeukaryotes to viruses 94%
Similar papers in this journal
- metaVaR: introducing metavariant species models for reference-free metagenomic-based population genomics 94%
- Precision long-read metagenomics sequencing for food safety by detection and assembly of Shiga toxin-producing Escherichia coli in irrigation water 94%
- centriflaken: an automated data analysis pipeline for assembly and in silico analyses of foodborne pathogens from metagenomic samples 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.