NanoVI: a Bayesian variational inference Nextflow pipelinefor species-level taxonomic classification from full-length16S rRNA Nanopore reads
Curiqueo, C.; Fuentes-Santander, F.; Ugalde, J. A.
Show abstract
SummaryNanoVI is a Nextflow pipeline for species-level taxonomic classification of full-length 16S rRNA Oxford Nanopore reads. Unlike existing tools that rely on expectation-maximization (EM) algorithms, NanoVI employs Bayesian variational inference with a Dirichlet-Categorical conjugate model, yielding abundance estimates accompanied by Bayesian 95% credible intervals that quantify estimation uncertainty, along with automatic shrinkage that suppresses spurious taxa. NanoVI integrates the Genome Taxonomy Database (GTDB) r226, providing phylogenetically consistent taxonomy while maintaining compatibility with NCBI-style databases. Benchmarked against a standardized mock community, NanoVI achieves species-detection metrics comparable to Emu, with 25-62% lower execution time and fewer false-positive assignments. Validation on 20 clinical vaginal microbiome samples confirms reproducibility against previously published Emu-based analyses. Availability and implementationNanoVI is implemented in Nextflow DSL2 with Docker containerization and is freely available at https://github.com/microbialds/NanoVI under an open-source license.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SCNIC: Sparse Correlation Network Investigation for Compositional Data 94%
- Unlocking River Biofilm Microbial Diversity: A Comparative Analysis of Sequencing Technologies 94%
- Obtaining deeper insights into microbiome diversity using a simple method to block host and non-targets in amplicon sequencing 94%
Similar papers in this journal
- Microbiome differential abundance methods produce disturbingly different results across 38 datasets 95%
- Integration of taxonomic signals from MAGs and contigs improves read annotation and taxonomic profiling of metagenomes 95%
- Genomic GC bias correction improves species abundance estimation from metagenomic data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.