Curating 16S rRNA databases enhances taxonomic accuracy and computational efficiency in microbial profiling
Baghbanzadeh, M.; Mahangade, V.; Crandall, K. A.; Rahnavard, A.
Show abstract
The 16S rRNA gene serves as the gold standard molecular marker for microbial profiling, yet taxonomic assignment accuracy depends critically on reference database quality. Substantial heterogeneity exists among databases in sequence coverage, curation standards, and taxonomic nomenclature, leading to conflicting taxonomic assignments. Despite previous comparisons highlighting performance differences, the impact of database preprocessing, including sequence cleaning and redundancy removal, on taxonomic classification remains understudied. To improve database quality, we implemented novel cleaning approaches to remove nested sequences, duplicate sequences, and correct missing taxonomic nomenclature. We compared four major 16S databases (SILVA, Greengenes2, RefSeq, and MIMt) using 69 mock communities and the DADA2 analysis pipeline for microbial genus-level profiling. Database size was significantly reduced after cleaning: SILVA reduced from 452,055 to 291,733 sequences, Greengenes2 from 337,506 to 277,982 sequences, MIMt from 48,749 to 34,734 sequences, and RefSeq from 27,376 to 25,970 sequences. Greengenes2, MIMt, and RefSeq, which exhibited comparable performance, consistently outperformed SILVA in recall, precision, and abundance estimation accuracy across all sample types. Cleaning SILVA improved computational efficiency by up to 50% while maintaining classification performance. We provide a benchmarking framework with the cleaned databases as resources for accurate 16S rRNA analysis and profiling.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Marine picoplankton metagenomes from eleven vertical profiles obtained by the Malaspina Expedition in the tropical and subtropical oceans 94%
- Highly accurate long-read HiFi sequencing data for five complex genomes 94%
- Characterizing organisms from three domains of life with universal primers from throughout the global ocean 94%
Similar papers in this journal
- Addressing the dynamic nature of reference data: a new nt database for robust metagenomic classification 96%
- GSR-DB: a manually curated and optimised taxonomical database for 16S rRNA amplicon analysis 95%
- Ecogenomics of groundwater viruses suggests niche differentiation linked to specific environmental tolerance 95%
Similar papers in this journal
- rRNA Operon Improves Species-Level Classification of Bacteria and Microbial Community Analysis Compared to 16S rRNA 95%
- Evaluating de novo assembly and binning strategies for time-series drinking water metagenomes. 94%
- Machine-learning based detection of adventitious microbes in T-cell therapy cultures using long read sequencing 93%
Similar papers in this journal
- Interactive Analysis of Biosurfactants in Fruit-Waste Fermentation Samples using BioSurfDB and MEGAN 94%
- Comprehensive Comparative Genomics Reveals Over 50 Phyla of Free-living and Pathogenic Bacteria are Associated with Diverse Members of the Amoebozoa 94%
- Robust bacterial co-occurence community structures are independent of r- and K-selection history 94%
Similar papers in this journal
- High-throughput DNA extraction and cost-effective miniaturized metagenome and amplicon library preparation of soil samples for DNA sequencing 94%
- CoSMIC - A hybrid approach for large-scale, high-resolution microbial profiling of novel niches 94%
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.