Parapipe: a Pipeline for Handling Parasite NGS Datasets and its Application to Cryptosporidium
Morris, A. V.; Robinson, G.; Chalmers, R.; Caccio, S.; Connor, T. R.
Show abstract
0.0Cryptosporidium, a protozoan parasite of significant public health concern, is responsible for severe diarrheal diseases, particularly in immunocompromised individuals and young children in resource-limited settings. Analysis of whole genome next generation sequencing (NGS) data is a critical next step in improving our understanding of Cryptosporidium epidemiology, transmission dynamics, and genetic diversity. However, effective analysis of NGS data in a public health context necessitates the development of robust, validated bioinformatics tools. Here, we present Parapipe, a modular ISO accreditable bioinformatics pipeline designed for high-throughput processing and analysis of Cryptosporidium NGS datasets. Built using Nextflow DSL2 and containerized with Singularity, Parapipe is portable, scalable, and capable of end-to-end analyses, including quality control, variant calling, multiplicity of infection (MOI) investigations and phylogenomic clustering analysis. Using both simulated and real-world datasets, we demonstrate Parapipes ability to resolve genetic heterogeneity, identify mixed infections, and generate high-resolution phylogenomic insights. Here, we use it to carry out a comparison between whole genome single nucleotide polymorphism (wgSNP) typing and the conventionally used gp60 molecular typing scheme. Compared to existing pipelines, Parapipe uniquely integrates MOI analysis, enabling the differentiation of mixed infections and supporting epidemiological investigations. Parapipes design facilitates integration with geographic, demographic, epidemiological and environmental data, enhancing its utility for tracking transmission pathways and outbreak sources. Parapipe represents a significant advance in utilising genomics for public health surveillance of Cryptosporidium, offering a streamlined and reproducible framework for analysis with potential application to other pathogenic protozoa. By automating complex workflows and enabling detailed genomic characterization, Parapipe provides a valuable tool for public health agencies and researchers, supporting efforts to mitigate the global burden of cryptosporidiosis.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Natrix: A Snakemake-based workflow for processing, clustering, and taxonomically assigning amplicon sequencing reads 94%
- HaVoC, a bioinformatic pipeline for reference-based consensus assembly and lineage assignment for SARS-CoV-2 sequences 94%
- HISS: Snakemake-based workflows for performing SMRT-RenSeq assembly, AgRenSeq and dRenSeq for the discovery of novel plant disease resistance genes. 94%
Similar papers in this journal
- Integrated population clustering and genomic epidemiology with PopPIPE 95%
- Bakta: Rapid & standardized annotation of bacterial genomes via alignment-free sequence identification 93%
- Development of a nextflow bioinformatics pipeline for the detection of SARS-CoV-2 co-infection cases from genomic surveillance in the Philippines 93%
Similar papers in this journal
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 94%
- AutoPhy: Automated phylogenetic identification of novel protein subfamilies 93%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 93%
Similar papers in this journal
- AMRomics: a scalable workflow to analyze large microbial genome collection 94%
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 94%
- SARS-CoV-2 surveillance in Italy through phylogenomic inferences based on Hamming distances derived from functional annotations of SNPs, MNPs and InDels 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.