MetaCerberus: distributed highly parallelized scalable HMM-based implementation for robust functional annotation across the tree of life
Figueroa, J. L.; Dhungel, E.; Brouwer, C. R.; White, R. A.
10.1101/2023.08.10.552700 bioRxivShow abstract
SummaryMetaCerberus is an exclusive HMM/HMMER-based tool that is massively parallel, on low memory, and provides rapid scalable annotation for functional gene inference across genomes to metacommunities. It provides robust enumeration of functional genes and pathways across many current public databases including KEGG (KO), COGs, CAZy, FOAM, and viral specific databases (i.e., VOGs and PHROGs). In a direct comparison, MetaCerberus was twice as fast as EggNOG-Mapper, and produced better annotation of viruses, phages, and archaeal viruses than DRAM, PROKKA, or InterProScan. MetaCerberus annotates more KOs across domains when compared to DRAM, with a 186x smaller database and a third less memory. MetaCerberus is fully integrated with differential statistical tools (i.e., DESeq2 and edgeR), pathway enrichment (GAGE R), and Pathview R for quantitative elucidation of metabolic pathways. MetaCerberus implements the key to unlocking the biosphere across the tree of life at scale. Availability and implementationMetaCerberus is written in Python and distributed under a BSD-3 license. The source code of MetaCerberus is freely available at https://github.com/raw-lab/metacerberus. Written in python 3 for both Linux and Mac OS X. MetaCerberus can also be easily installed using mamba create -n metacerberus -c bioconda -c conda-forge metacerberus
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Unveiling the Microbial Realm with VEBA 2.0: A modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic, and viral multi-omics from either short- or long-read sequencing 97%
- Automating microbial taxonomy workflows with PHANTASM: PHylogenomic ANalyses for the TAxonomy and Systematics of Microbes 96%
- BGCFlow: Systematic pangenome workflow for the analysis of biosynthetic gene clusters across large genomic datasets 96%
Similar papers in this journal
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 95%
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 95%
- Extraction of near-complete genomes from metagenomic samples: a new service in PATRIC 94%
Similar papers in this journal
- ChiMera: An easy to use pipeline for Bacterial Genome Based Metabolic Network Reconstruction, Evaluation and Visualization 95%
- PoMeLo: a systematic computational approach to predicting metabolic loss in pathogen genomes 94%
- Ribovore: ribosomal RNA sequence analysis for GenBank submissions and database curation 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.