Carbohydrate-active enzyme annotation in microbiomes using dbCAN
Zheng, J.; Huang, L.; Yi, H.; Yan, Y.; Zhang, X.; Akresi, J.; Yin, Y.
Show abstract
CAZymes or carbohydrate-active enzymes are critically important for human gut health, lignocellulose degradation, global carbon recycling, soil health, and plant disease. We developed dbCAN as a web server in 2012 and actively maintain it for automated CAZyme annotation. Considering data privacy and scalability, we provide run_dbcan as a standalone software package since 2018 to allow users perform more secure and scalable CAZyme annotation on their local servers. Here, we offer a comprehensive computational protocol on automated CAZyme annotation of microbiome sequencing data, covering everything from short read pre-processing to data visualization of CAZyme and glycan substrate occurrence and abundance in multiple samples. Using a real-world metagenomic sequencing dataset, this protocol describes commands for dataset and software preparation, metagenome assembly, gene prediction, CAZyme prediction, CAZyme gene cluster (CGC) prediction, glycan substrate prediction, and data visualization. The expected results include publication-quality plots for the abundance of CAZymes, CGCs, and substrates from multiple CAZyme annotation routes (individual sample assembly, co-assembly, and assembly-free). For the individual sample assembly route, this protocol takes [~]33h on a Linux computer with 40 CPUs, while other routes will be faster. This protocol does not require programming experience from users, but it does assume a familiarity with the Linux command-line interface and the ability to run Python scripts in the terminal. The target audience includes the tens of thousands of microbiome researchers who routinely use our web server. This protocol will encourage them to perform more secure, rapid, and scalable CAZyme annotation on their local computer servers.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Tutorial: Assessing metagenomics software with the CAMI benchmarking toolkit 95%
- Inferring cellular and molecular processes in single-cell data with non-negative matrix factorization using Python, R, and GenePattern Notebook implementations of CoGAPS 94%
- kallisto, bustools, and kb-python for quantifying bulk, single-cell, and single-nucleus RNA-seq 93%
Similar papers in this journal
- pISA-tree - a data management framework for life science research projects using a standardised directory tree 95%
- Highly accurate long-read HiFi sequencing data for five complex genomes 94%
- Marine picoplankton metagenomes from eleven vertical profiles obtained by the Malaspina Expedition in the tropical and subtropical oceans 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.