Cascabel: a flexible, scalable and easy-to-use amplicon sequence data analysis pipeline
Abdala Asbun, A.; Besseling, M. A.; Balzano, S.; van Bleijswijk, J.; Witte, H.; Villanueva, L.; Engelmann, J. C.
Show abstract
Marker gene sequencing of the rRNA operon (16S, 18S, ITS) or cytochrome c oxidase I (CO1) is a popular means to assess microbial communities of the environment, microbiomes associated with plants and animals, as well as communities of multicellular organisms via environmental DNA sequencing. Since this technique is based on sequencing a single gene rather than the entire genome, the number of reads needed per sample is lower than that required for metagenome sequencing, making marker gene sequencing affordable to nearly any laboratory. Despite the relative ease and cost-efficiency of data generation, analyzing the resulting sequence data requires computational skills that may go beyond the standard repertoire of a current molecular biologist/ecologist. We have developed Cascabel, a flexible and easy-to-use amplicon sequence data analysis pipeline, which uses Snakemake and a combination of existing and newly developed solutions for its computational steps. Cascabel takes the raw data as input and delivers a table of operational taxonomic units (OTUs) and a representative sequence tree. Our pipeline allows customizing the analyses by offering several choices for most of the steps, for example different OTU generating methods. The pipeline can make use of multiple computing nodes and scales from personal computers to computing servers. The analyses and results are fully reproducible and documented in an HTML and optional pdf report. Cascabel is freely available at Github: https://github.com/AlejandroAb/CASCABEL and licensed under GNU GPLv3.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natrix: A Snakemake-based workflow for processing, clustering, and taxonomically assigning amplicon sequencing reads 96%
- ATLAS: a Snakemake workflow for assembly, annotation, and genomic binning of metagenome sequence data 95%
- Ribovore: ribosomal RNA sequence analysis for GenBank submissions and database curation 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.