Emu: Species-Level Microbial Community Profiling for Full-Length Nanopore 16S Reads
Curry, K. D.; Wang, Q.; Nute, M. G.; Tyshaieva, A.; Reeves, E.; Soriano, S.; Graeber, E.; Finzer, P.; Mendling, W.; Wu, Q.; Savidge, T.; Villapol, S.; Dilthey, A.; Treangen, T. J.
Show abstract
16S rRNA based analysis is the established standard for elucidating microbial community composition. While short read 16S analyses are largely confined to genus-level resolution at best since only a portion of the gene is sequenced, full-length 16S sequences have the potential to provide species-level accuracy. However, existing taxonomic identification algorithms are not optimized for the increased read length and error rate of long-read data. Here we present Emu, a novel approach that employs an expectation-maximization (EM) algorithm to generate taxonomic abundance profiles from full-length 16S rRNA reads. Results produced from one simulated data set and two mock communities prove Emu capable of accurate microbial community profiling while obtaining fewer false positives and false negatives than alternative methods. Additionally, we illustrate a real-world application of our new software by comparing clinical sample composition estimates generated by an established whole-genome shotgun sequencing workflow to those returned by full-length 16S sequences processed with Emu.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MetaPro: A scalable and reproducible data processing and analysis pipeline for metatranscriptomic investigation of microbial communities 97%
- Ultra-accurate Microbial Amplicon Sequencing with Synthetic Long Reads 96%
- Identifying unmeasured heterogeneity in microbiome data via quantile thresholding (QuanT) 96%
Similar papers in this journal
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 96%
- BiG-MAP: an automated pipeline to profile metabolic gene cluster abundance and expression in microbiomes 95%
- parafac4microbiome: Exploratory analysis of longitudinal microbiome data using Parallel Factor Analysis 95%
Similar papers in this journal
- SCNIC: Sparse Correlation Network Investigation for Compositional Data 96%
- Obtaining deeper insights into microbiome diversity using a simple method to block host and non-targets in amplicon sequencing 96%
- A flexible pipeline combining clustering and correction tools for prokaryotic and eukaryotic metabarcoding 95%
Similar papers in this journal
Similar papers in this journal
- Robust bacterial co-occurence community structures are independent of r- and K-selection history 96%
- Comparison of the effectiveness of different normalization methods for metagenomic cross-study phenotype prediction under heterogeneity 95%
- Meta-analysis of Microbiome Association Networks Reveal Patterns of Dysbiosis in Diseased Microbiomes 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.