Back

Full 16S and 23S rRNA gene-based,strain-level resolution of the microbiota of a mock bacterial microbiome using ONT nanopore sequencing

Woodruff, C. J.; Zhang, Z.; Speed, T. P.

2024-02-09 bioinformatics
10.1101/2024.02.07.579414 bioRxiv
Show abstract

PurposeThe feasibility of achieving near-strain taxonomic resolution and reliable microbiota strain-level abundance estimation using amplicon-based nanopore sequencing was investigated. MethodsDenoising was applied to separate 16S and 23S genes extracted from a metagenomic dataset generated by Sereika et al. [9] using nanopore sequencing. Denoising used Kumar et al.s [8] Robust Amplicon Denoising (RAD) and generated amplicon sequence variants (ASVs). The Sereika et al. data set was generated from the Zymo [10] D6322 7-bacterial species mock microbiome. A sub-sampling procedure generated additional datasets that allowed sensitivity assessment over multiple orders of relative abundance. Alignment to bespoke 16S and 23S rRNA gene databases, provided both identification and abundance information data for sub-species analyses. ResultsSub-species identification was clearly achieved, both with the "even" D6322 dataset, and with the 4 sub-sampled datasets in which 3-orders of magnitude relative abundances were present. Multiple ASVs of length approximately 2500 bases were generated which gave perfect alignments to reference 23S rRNA genes. Similarly for the approximately 1500 base long 16S rRNA genes. Despite strain ambiguity for some species, the strain of each species known to be present was identified in all but 1 (of 70) cases examined for all species. A process to merge the 16S and 23S rRNA genes results reduced ambiguity, allowing better sub-species resolution, and better species, and sub-species, relative cellular abundance estimates. ConclusionPrincipled methods for dealing with strain-level ambiguity in microbiota analysis have been developed that are generally applicable. Their application to nanopore sequencing with the current state of Oxford Nanopore Technology, together with state-of-the-art denoising, allows near strain resolution of bacterial microbiota identity, and improved estimates of species, and sub-species, relative abundances.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.