sangerFlow, a Sanger sequencing-based bioinformatics pipeline for pests and pathogens identification
Prodhan, M. A.; Power, M.; Kehoe, M.
Show abstract
Sequencing of a Polymerase Chain Reaction product (amplicon) is called amplicon sequencing. Amplicon sequencing allows for reliable identification of an organism by amplifying, sequencing, and analysing a single conserved marker gene or DNA barcode. As this approach generally involves a single gene, it is a light-weight protocol compared to multi-locus or whole genome sequencing for diagnostic purposes; yet considerably reliable. Therefore, Sanger-based high-quality amplicon sequencing is widely deployed for species identification and high-throughput biosecurity surveillance. However, keeping up with the data analysis in a large-scale surveillance or diagnostic settings could be a limiting factor because it involves manual quality control of the raw sequencing data, alignment of the forward and reverse reads, and finally web-based Blastn search of all the amplicons. Here, we present a bioinformatics pipeline that automates the entire analysis. As a result, the pipeline is scalable with high-volume of samples and reproducible. Furthermore, the pipeline leverages the modern open-source Nextflow and Singularity concept, thus it does not require software installation except Nextflow and Singularity, software subscription, or programming expertise from the end users making it widely adaptable. Availability and implementationsangerFlow source code and documentation are freely available for download at GitHub, implemented in Nextflow and Singularity.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natrix: A Snakemake-based workflow for processing, clustering, and taxonomically assigning amplicon sequencing reads 94%
- Comprehensive benchmarking of metagenomic classification tools for long-read sequencing data 94%
- A comparison of three programming languages for a full-fledged next-generation sequencing tool 94%
Similar papers in this journal
- Reliable variant calling during runtime of Illumina sequencing 94%
- Detecting SARS-CoV-2 lineages and mutational load in municipal wastewater; a use-case in the metropolitan area of Thessaloniki, Greece 94%
- Design of Specific Primer Set for Detection of B.1.1.7 SARS-CoV-2 Variant using Deep Learning 93%
Similar papers in this journal
- Runcer-Necromancer: A Method To Rescue Data From An Interrupted Run On MGISEQ-2000 93%
- NAD: Noise-augmented direct sequencing of target nucleic acids by augmenting with noise and selective sampling 92%
- RNAmining: A machine learning stand-alone and web server tool for RNA coding potential prediction 91%
Similar papers in this journal
- CompareM2 is a genomes-to-report pipeline for comparing microbial genomes 95%
- ganon: precise metagenomics classification against large and up-to-date sets of reference sequences 94%
- MLDSP-GUI: An alignment-free standalone tool with an interactive graphical user interface for DNA sequence comparison and analysis 93%
Similar papers in this journal
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 94%
- Haxe as a Swiss knife for bioinformatic applications: the SeqPHASE case story 93%
- Comparing full variation profile analysis with the conventional consensus method in SARS-CoV-2 phylogeny 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.