Taxonize-gb: A tool for filtering GenBank non-redundant databases based on taxonomy
Sarhan, M. S.; Filosi, M.; Maixner, F.; Fuchsberger, C.
Show abstract
Analyzing taxonomic diversity and identification in diverse ecological samples has become a crucial routine in various research and industrial fields. While DNA barcoding marker-gene approaches were once prevalent, the decreasing costs of next-generation sequencing have made metagenomic shotgun sequencing more popular and feasible. In contrast to DNA-barcoding, metagenomic shotgun sequencing offers possibilities for in-depth characterization of structural and functional diversity. However, analysis of such data is still considered a hurdle due to absence of taxa-specific databases. Here we present taxonize-gb, a command-line software tool to extract GenBank non-redundant nucleotide and protein databases, related to one or more input taxonomy identifier. Our tool allows the creation of taxa-specific reference databases tailored to specific research questions, which reduces search times and therefore represents a practical solution for researchers analyzing large metagenomic data on regular basis. Taxonize-gb is an open-source command-line Python-based tool freely available for installation at https://pypi.org/project/taxonize-gb/ and on GitHub https://github.com/msabrysarhan/taxonize_genbank. It is released under Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0).
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AutoPhy: Automated phylogenetic identification of novel protein subfamilies 94%
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 93%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.