Benchmarking the impact of reference genome selection on taxonomic profiling accuracy
van Bemmelen, J.; Nika, I.; Baaijens, J. A.
Show abstract
BackgroundOver the past decades, genome databases have expanded exponentially, often incorporating highly similar genomes at the same taxonomic level. This redundancy can hinder taxonomic classification, leading to difficulties distinguishing between closely related sequences and increasing computational demands. While some novel taxonomic classification tools address this redundancy by selecting a subset of genomes as references, insights regarding the impact of different reference genome selection methods across taxonomic classification tools are lacking. ResultsWe systematically evaluate genome selection and dereplication methods on bacterial and viral datasets using simulated metagenomic samples. For bacterial species-level profiling, incorporating all available genomes generally yields the highest accuracy, while having a limited impact on computational resource usage. In contrast, for highly similar bacterial strain-level and SARS-CoV-2 lineage-level datasets we find that selection significantly improves abundance estimation accuracy. Incorporating location-based metadata further enhances viral profiling performance by prioritizing locally relevant genomes. Across viral experiments, smaller reference sets significantly reduce memory and runtime requirements during both indexing and profiling, although this comes at an additional pre-processing cost. ConclusionsReference genome selection influences both accuracy and computational efficiency in taxonomic profiling, but its benefits seem context- and resolution-dependent. Our results demonstrate that reference set design does not have a one-size-fits-all solution, and that selection strategies should be adapted based on the biological and computational setting.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 97%
- Read-SpaM: assembly-free and alignment-free comparison of bacterial genomes with low sequencing coverage 97%
- Sketching and sampling approaches for fast and accurate long read classification 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.