Back

Benchmarking the impact of reference genome selection on taxonomic profiling accuracy

van Bemmelen, J.; Nika, I.; Baaijens, J. A.

2025-02-08 bioinformatics
10.1101/2025.02.07.637076 bioRxiv
Show abstract

BackgroundOver the past decades, genome databases have expanded exponentially, often incorporating highly similar genomes at the same taxonomic level. This redundancy can hinder taxonomic classification, leading to difficulties distinguishing between closely related sequences and increasing computational demands. While some novel taxonomic classification tools address this redundancy by selecting a subset of genomes as references, insights regarding the impact of different reference genome selection methods across taxonomic classification tools are lacking. ResultsWe systematically evaluate genome selection and dereplication methods on bacterial and viral datasets using simulated metagenomic samples. For bacterial species-level profiling, incorporating all available genomes generally yields the highest accuracy, while having a limited impact on computational resource usage. In contrast, for highly similar bacterial strain-level and SARS-CoV-2 lineage-level datasets we find that selection significantly improves abundance estimation accuracy. Incorporating location-based metadata further enhances viral profiling performance by prioritizing locally relevant genomes. Across viral experiments, smaller reference sets significantly reduce memory and runtime requirements during both indexing and profiling, although this comes at an additional pre-processing cost. ConclusionsReference genome selection influences both accuracy and computational efficiency in taxonomic profiling, but its benefits seem context- and resolution-dependent. Our results demonstrate that reference set design does not have a one-size-fits-all solution, and that selection strategies should be adapted based on the biological and computational setting.

Published in BMC Genomics (predicted rank #13) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.