Back

Curating 16S rRNA databases enhances taxonomic accuracy and computational efficiency in microbial profiling

Baghbanzadeh, M.; Mahangade, V.; Crandall, K. A.; Rahnavard, A.

2025-11-05 bioinformatics
10.1101/2025.11.04.686545 bioRxiv
Show abstract

The 16S rRNA gene serves as the gold standard molecular marker for microbial profiling, yet taxonomic assignment accuracy depends critically on reference database quality. Substantial heterogeneity exists among databases in sequence coverage, curation standards, and taxonomic nomenclature, leading to conflicting taxonomic assignments. Despite previous comparisons highlighting performance differences, the impact of database preprocessing, including sequence cleaning and redundancy removal, on taxonomic classification remains understudied. To improve database quality, we implemented novel cleaning approaches to remove nested sequences, duplicate sequences, and correct missing taxonomic nomenclature. We compared four major 16S databases (SILVA, Greengenes2, RefSeq, and MIMt) using 69 mock communities and the DADA2 analysis pipeline for microbial genus-level profiling. Database size was significantly reduced after cleaning: SILVA reduced from 452,055 to 291,733 sequences, Greengenes2 from 337,506 to 277,982 sequences, MIMt from 48,749 to 34,734 sequences, and RefSeq from 27,376 to 25,970 sequences. Greengenes2, MIMt, and RefSeq, which exhibited comparable performance, consistently outperformed SILVA in recall, precision, and abundance estimation accuracy across all sample types. Cleaning SILVA improved computational efficiency by up to 50% while maintaining classification performance. We provide a benchmarking framework with the cleaned databases as resources for accurate 16S rRNA analysis and profiling.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.