Taxallnomy: Closing gaps in the NCBI Taxonomy
Sakamoto, T.; Ortega, J. M.
Show abstract
NCBI Taxonomy is the main taxonomic source for several bioinformatics tools and databases since all organisms with sequence accessions deposited on INSDC are organized in its hierarchical structure. Despite the extensive use and application of this data source, taking advantage of its taxonomic tree could be challenging because (1) some taxonomic ranks are missing in some lineages and (2) some nodes in the tree do not have a taxonomic rank assigned (referred to as "no rank"). To address this issue, we developed an algorithm that takes the tree structure from NCBI Taxonomy and generates a hierarchically complete taxonomic tree. The procedures performed by the algorithm consist of attempting to assign a taxonomic rank to "no rank" nodes and of creating/deleting nodes throughout the tree. The algorithm also creates a name for the new nodes by borrowing the names from its ranked child or, if there is no child, from its ranked parent node. The new hierarchical structure was named taxallnomy and it contains 33 hierarchical levels corresponding to the 33 taxonomic ranks currently used in the NCBI Taxonomy database. From taxallnomy, users can obtain the complete taxonomic lineage with 33 nodes of all taxa available in the NCBI Taxonomy database. Taxallnomy is applicable to several bioinformatics analyses that depend on NCBI Taxonomy data. In this work, we demonstrate its applicability by embedding taxonomic information of a specified rank into a phylogenetic tree; and by making metagenomics profiles. Taxallnomy algorithm was written in PERL and all its resources are available at bioinfo.icb.ufmg.br/taxallnomy. Database URL: http://bioinfo.icb.ufmg.br/taxallnomy
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Estimating body volumes and surface areas of animalsfrom cross-sections 93%
- Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors 92%
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 92%
Similar papers in this journal
- Machine learning based imputation techniques for estimating phylogenetic trees from incomplete distance matrices 94%
- MeShClust v3.0: High-quality clustering of DNA sequences using the mean shift algorithm and alignment-free identity scores 93%
- Comprehensive analysis of 111 Pleuronectiformes mitochondrial genomes: insights into structure, conservation, variation and evolution 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.