Incorporation of data from multiple hypervariable regions when analyzing bacterial 16S rRNA sequencing data
Jones, C. B.; White, J. R.; Ernst, S. E.; Sfanos, K. S.; Peiffer, L. B.
Show abstract
Short read 16S rRNA amplicon sequencing is a common technique used in microbiome research. However, inaccuracies in estimated bacterial community composition can occur due to amplification bias of the targeted hypervariable region. A potential solution is to sequence and assess multiple hypervariable regions in tandem, yet there is currently no consensus as to the appropriate method for analyzing this data. Additionally, there are many sequence analysis resources for data produced from the Illumina platform, but fewer open-source options available for data from the Ion Torrent platform. Herein, we present an analysis pipeline using an open-source analysis platform that integrates data from multiple hypervariable regions and is compatible with data produced from the Ion Torrent platform. We used the ThermoFisher Ion 16S Metagenomics Kit and a mock community of 20 bacterial strains to assess taxonomic classification of amplicons from 6 separate hypervariable regions (V2, V3, V4, V6-7, V8, V9) using our analysis pipeline. We report that different hypervariable regions have different specificities for taxonomic classification, which also had implications for global level analyses such as alpha and beta diversity. Finally, we utilize a generalized linear modeling approach to statistically integrate the results from multiple hypervariable regions and apply this methodology to data from a small clinical cohort. We conclude that scrutinizing sequencing results separately by hypervariable region provides a more granular view of the taxonomic classification achieved by each primer set as well as the concordance of results across hypervariable regions. However, the data across all hypervariable regions can be combined using generalized linear models to statistically evaluate overall differences in community structure and relatedness among sample groups.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- rRNA Operon Improves Species-Level Classification of Bacteria and Microbial Community Analysis Compared to 16S rRNA 98%
- Library Preparation and Sequencing Platform Introduce Bias in Metagenomic-Based Characterizations of Microbiomes 96%
- Nested PCR to optimize rpoB metabarcoding for low-concentration and host-associated bacterial DNA 96%
Similar papers in this journal
- GSR-DB: a manually curated and optimised taxonomical database for 16S rRNA amplicon analysis 97%
- Clade-specific long-read sequencing increases the accuracy and specificity of the gyrB phylogenetic marker gene 96%
- Two-target quantitative PCR to predict library composition for shallow shotgun sequencing 96%
Similar papers in this journal
- Finding the right fit: A comprehensive evaluation of short-read and long-read sequencing approaches to maximize the utility of clinical microbiome data 96%
- Benchmarking taxonomic classifiers with Illumina and Nanopore sequence data for clinical metagenomic diagnostic applications 95%
- Duplex real-time PCR assay for the simultaneous detection of Achromobacter xylosoxidans and Achromobacter spp. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.