AmpliconTyper tool for analysing ONT multiplex PCR data from environmental and other mixed sources
Spadar, A.; Mahindroo, J.; Troman, C.; Owusu, M.; Abu-Sarkodie, Y.; Owusu-Dabo, E.; Abraham, D.; Blossom, B.; Govindan, K.; Mohan, V. R.; Dyson, Z. A.; Grassly, N. C.; Holt, K.
Show abstract
Amplicon sequencing is a popular method for understanding the diversity of bacterial communities in mixed samples as exemplified by 16S rRNA metagenome sequencing. This approach has been extended into multiplex amplicon sequencing in which multiple targets are amplified in the same polymerase chain reaction (PCR). Multiple tools exist to process the sequencing data produced via the short-read Illumina platform, but there are fewer options for long-read Oxford Nanopore Technologies (ONT) sequencing, or for processing data from environmental surveillance or other sources with many different organisms. We have developed AmpliconTyper (v0.1.28, DOI: 10.5281/zenodo.14621928) for analysing multiplex amplicon sequencing data from environmental (e.g. wastewater) or similarly contaminated samples, generated using ONT devices. The software tool uses machine learning to classify sequencing reads into target and non-target organisms with very high specificity and sensitivity. The user can train models using public and/or user-generated data, which can subsequently be applied to analyse new data. The tool can also generate amplicon consensus sequences, as well as identify single nucleotide polymorphisms (SNPs) and report their genotype implications, such as association with lineages or antimicrobial resistance (AMR). The tool is freely available via Bioconda and GitHub (https://github.com/AntonS-bio/AmpliconTyper). AmpliconTyper allows robust identification of target organism reads in ONT sequenced environmental samples, and can identify user-specified lineage or AMR markers. Impact statementAmpliconTyper (v0.1.28) is a software package that enables users to analyse amplicon sequences generated by targeted amplification followed by ONT sequencing, for environmental or other similarly contaminated samples. The analysis includes mapping of reads to target amplicon sequences, classification of each sequenced read as either originating from a target or non-target organism, followed by identification of user-specified SNPs and generation of an interactive report summarising the findings. The strength of AmpliconTyper lies in its ability to train a machine learning model using public data to create sequencing read classification models tailored to a users application. AmpliconTyper is designed specifically to work with extremely noisy data that includes a large share of off-target amplification reads such as those encountered in environmental surveillance applications. Data SummaryFor the purpose of designing and testing AmpliconTyper we have used two datasets. The first consisted of 69 Salmonella enterica serovar. Typhi (S. Typhi) and 10,303 other Enterobacteriaceae WGS ONT nanopore sequencing libraries (Supp. Data 1) from NCBI Sequence Read Archive (SRA) (1). We used these data to evaluate the performance of different classifier models and to train a model for our use-case, i.e. amplicon-based detection of S. Typhi from environmental surveillance samples (2). In addition, to further evaluate the performance of AmpliconTyper in our use case, we have applied an amplicon sequencing protocol (2) to generate amplicon data for S. Typhi using two S. Typhi isolates (NCBI accessions SRR5949979 and SRR7165748) provided by Satheesh Nair (UKHSA) (3). We also used same protocol to the pooled sample of American Type Culture Collection (ATCC) consisting of S. Paratyphi A (ATCC 9150D), S. Paratyphi B (ATCC-BAA-1250D), S. Paratyphi C (ATCC-BAA-1715D), Aeromonas hydrophila (ATCC-7965D), Klebsiella pneumoniae (ATCC-BAA-1706D), and Citrobacter freundii (ATCC-8090D) chosen for their close relationship to S. Typhi. The test data for classification is available at https://github.com/AntonS-bio/AmpliconTyper/tree/main/test_data. Newly generated data was deposited in European Nucleotide Archive project PRJEB81565.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Unveiling the Microbial Realm with VEBA 2.0: A modular bioinformatics suite for end-to-end genome-resolved prokaryotic, (micro)eukaryotic, and viral multi-omics from either short- or long-read sequencing 95%
- Automating microbial taxonomy workflows with PHANTASM: PHylogenomic ANalyses for the TAxonomy and Systematics of Microbes 94%
- Linkage-based ortholog refinement in bacterial pangenomes with CLARC 94%
Similar papers in this journal
- Accurate and Reproducible Whole-Genome Genotyping for Bacterial Genomic Surveillance with Nanopore Sequencing Data 96%
- Identifying indels from WGS short reads of haploid genomes distinguishes variant-calling algorithms 95%
- The clinical utility of two high-throughput 16S rRNA gene sequencing workflows for taxonomic assignment of unidentifiable bacterial pathogens in MALDI-TOF MS 93%
Similar papers in this journal
- HighALPS: Ultra-High-Throughput Marker-Gene Amplicon Library Preparation and Sequencing on the Illumina NextSeq and NovaSeq Platforms 94%
- Species level resolution of female bladder microbiota from 16S rRNA amplicon sequencing 94%
- Two-target quantitative PCR to predict library composition for shallow shotgun sequencing 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.