The Metabarcoding Analysis Pipeline (MAP): Simple, accurate, and flexible metabarcoding
Prosser, S. W.; Bard, N. W.; Thompson, K. A.; Floyd, R. A.; Padhye, S.; Ozsahin, E.; Jafarpour, S.; Hebert, P. D. N.
Show abstract
Current metabarcoding pipelines are inflexible with respect to study design and are poorly suited to long-read sequence data. To address these limitations, we developed MAP, the Metabarcoding Analysis Pipeline, which is a sequence-to-answer workflow supporting the analysis of amplicons from highly multiplexed and replicated study designs. Although MAP can analyze amplicons of any length from any genetic marker, it includes several features tailored to long-read COI metabarcoding. MAP installs from a Docker container and requires only sequence data, a parameters file, and a reference library. It produces intuitive reports, enabling users to evaluate their data immediately after analysis. We validate MAP by showing that it generates biodiversity estimates that correspond closely to a ground-truth dataset of single-specimen DNA barcode data and by demonstrating that it outperforms alternative platforms for COI metabarcoding. MAP is free, open-source, and available from: https://github.com/cbg-innov/MAP.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Iroki: automatic customization and visualization of phylogenetic trees 93%
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 92%
- VirION2: a short- and long-read sequencing and informatics workflow to study the genomic diversity of viruses in nature 92%
Similar papers in this journal
- Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets 95%
- Reverse engineering environmental metatranscriptomes clarifies best practices for eukaryotic assembly 94%
- SQMtools: automated processing and visual analysis of 'omics data with R and anvi'o 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.