Optimizing microbial community metagenomic hybrid assembly correction and polishing approaches informed by reference-independent analyses
Smith, G. J.; van Alen, T.; van Kessel, M.; Luecker, S.
Show abstract
Hybrid metagenomic assembly, leveraging both long- and short-read sequencing technologies, of microbial communities is becoming an increasingly accessible approach, yet its widespread application faces several challenges. High-quality references may not be available for assembly accuracy comparisons common for benchmarking, and certain aspects of hybrid assembly may require dataset-dependent, empirically-guided optimization rather than application of a uniform approach. In this study, several simple, reference-free characteristics - gene lengths and read recruitment - were analyzed as reliable proxies of assembly quality to guide hybrid assembly optimization. These characteristics were further explored in relation to reference-dependent genome- and gene-centric analyses that are common for microbial community metagenomic studies. Here, two laboratory-scale bioreactors were sequenced with short and long read platforms, and assembled with commonly used software packages. Following long read assembly, long read correction and short read polishing were iterated to resolve errors. Each iteration in this process was shown so have a substantial effect on gene- and genome-centric community composition. Simple, reference-free assembly characteristics, specifically changes in gene fragmentation and short read recruitment, explored throughout this process replicated patterns of more advanced analyses seen in published comparative studies, and therefore are suitable proxies for hybrid metagenome assembly accuracy to save computational resources. Hybrid metagenomic sequencing approaches will likely remain relevant due to the low costs of short read sequencing, therefore it is imperative that users are equipped to estimate assembly accuracy prior to downstream gene- and genome-centric analyses.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Longitudinal, Multi-platform Metagenomics Yields a High-quality Genomic Catalog and Guides an In Vitro Model for Cheese Communities 96%
- Preparation of functional metagenomic libraries from low biomass samples using METa assembly and their application to capture antibiotic resistance genes 95%
- A reproducible and tunable synthetic soil microbial community provides new insights into microbial ecology 95%
Similar papers in this journal
- BinaRena: a dedicated interactive platform for human-guided exploration and binning of metagenomes 96%
- Increased Replication Rates of Dissimilatory Nitrogen-Reducing Bacteria Leads to Decreased Anammox Reactor Performance 96%
- MetaPro: A scalable and reproducible data processing and analysis pipeline for metatranscriptomic investigation of microbial communities 96%
Similar papers in this journal
- MiDAS 5: Global diversity of bacteria and archaea in anaerobic digesters 96%
- Automated strain separation in low-complexity metagenomes using long reads 95%
- MiDAS 4: A global catalogue of full-length 16S rRNA gene sequences and taxonomy for studies of bacterial communities in wastewater treatment plants 95%
Similar papers in this journal
- Robust bacterial co-occurence community structures are independent of r- and K-selection history 94%
- Adaptations of endolithic communities to abrupt environmental changes in a hyper-arid desert 93%
- Unravelling microalgal-bacterial interactions in aquatic ecosystems through 16S rRNA gene-based co-occurrence networks 93%
Similar papers in this journal
- Recovery of small plasmid sequences via Oxford Nanopore sequencing 96%
- From defaults to databases: parameter and database choice dramatically impact the performance of metagenomic taxonomic classification tools 95%
- Finding the right fit: A comprehensive evaluation of short-read and long-read sequencing approaches to maximize the utility of clinical microbiome data 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.