Back

NanoVI: a Bayesian variational inference Nextflow pipelinefor species-level taxonomic classification from full-length16S rRNA Nanopore reads

Curiqueo, C.; Fuentes-Santander, F.; Ugalde, J. A.

2026-03-10 bioinformatics
10.64898/2026.03.07.710315 bioRxiv
Show abstract

SummaryNanoVI is a Nextflow pipeline for species-level taxonomic classification of full-length 16S rRNA Oxford Nanopore reads. Unlike existing tools that rely on expectation-maximization (EM) algorithms, NanoVI employs Bayesian variational inference with a Dirichlet-Categorical conjugate model, yielding abundance estimates accompanied by Bayesian 95% credible intervals that quantify estimation uncertainty, along with automatic shrinkage that suppresses spurious taxa. NanoVI integrates the Genome Taxonomy Database (GTDB) r226, providing phylogenetically consistent taxonomy while maintaining compatibility with NCBI-style databases. Benchmarked against a standardized mock community, NanoVI achieves species-detection metrics comparable to Emu, with 25-62% lower execution time and fewer false-positive assignments. Validation on 20 clinical vaginal microbiome samples confirms reproducibility against previously published Emu-based analyses. Availability and implementationNanoVI is implemented in Nextflow DSL2 with Docker containerization and is freely available at https://github.com/microbialds/NanoVI under an open-source license.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.