Back

Arche: A Hierarchical, Easy To Download, Functional-Optimized Annotator For Microbial Meta(Genomes)

Alonso-Reyes, D. G.; Albarracin, V. H.

2022-11-29 bioinformatics
10.1101/2022.11.28.518280 bioRxiv
Show abstract

The growing amount of genomic data has prompted a need for less demanding and user friendly functional annotators. At the present, its hard to find a pipeline for the annotation of multiple functional data, such as both enzyme commission numbers (E.C.) and orthologous identifiers (KEGG and eggNOG), protein names, gene names, alternative names, and descriptions. In this work, we provide a new solution which combines different algorithms (BLAST, DIAMOND, HMMER3) and databases (UniprotKB, KOfam, NCBIFAMs, TIGRFAMs, and PFAM), and also overcome data download challenges. The software framework, Arche, herein demonstrated competitive results over Escherichia coli K-12 genome, the metagenome-assembled genome of Ferrovum myxofaciens S2.4, and a freshwater metagenome when compared to other annotators. Finally, Arche provides an analysis pipeline that can accommodate advanced tools in a unique order, creating several advantages regarding to other commonly used annotators.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.