STECode: an automated virulence barcode generator to aid clinical and public health risk assessment of Shiga toxin-producing Escherichia coli
Sim, E. M.; Fong, W.; Suster, C.; Agius, J. E.; Chandra, S.; Suliman, B.; Wang, Q.; Ngo, C.; Finemore, C.; Chen, S. C.-A.; Basile, K.; Sintchenko, V.
Show abstract
Shiga Toxin (Stx) producing Escherichia coli (STEC) is a subset of pathogenic E. coli that can produce two types of Stx, Stx1 and Stx2, which can be further subtyped into four and 15 subtypes respectively. Not all subtypes, however, are equal in virulence potential, and the risk of severe disease including haemolytic uraemic syndrome has been linked to certain Stx2 subtypes e.g. Stx2a, Stx2d, highlighting the importance to survey stx subtypes. Previously, we developed a STEC virulence barcode to capture pertinent information on virulence genes to infer pathogenic potential. However, the process required multiple manual curation steps to determine the barcode. Here we introduce STECode, a bioinformatic tool to automate the STEC virulence barcode generation from sequencing reads or genomic assemblies. The development, and validation of STECode is described using a set of publicly available completed STEC genomes, along with their corresponding short reads. STECode was applied to interrogate the virulence landscape and molecular epidemiology of human STEC isolated during the period of the international border closures related to COVID-19 in the state of New South Wales, Australia. Impact statementWhole genome sequencing has been used to great effect in the genomic surveillance of STEC for public health purposes via the tracking of outbreaks. With STECode, we present a method to generate a STEC virulence barcode which captures pertinent subtyping information, useful for genomic inference of pathogenic potential. A key blind spot generated in short-read sequencing is the inability to detect the presence of multiple, isogenic stx copies in STEC. STECode mitigates this by inferring and reporting on the possibility of this occurrence. We envisage that this tool will value-add current genomic surveillance workflows through the ability to infer pathogenic potential.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- K-mer based prediction of Clostridioides difficile relatedness and ribotypes 96%
- A high quality reference genome for the fish pathogen Streptococcus iniae 95%
- Phylogenomic analysis of Clostridium perfringens identifies isogenic strains in gastroenteritis outbreaks, and novel virulence-related features 95%
Similar papers in this journal
- Achromobacter xylosoxidans isolates exhibit genome diversity, variable virulence, high levels of antibiotic resistance and potential intrahost evolution. 95%
- Genes influencing phage host range in Staphylococcus aureus on a species-wide scale 94%
- Diversity and prevalence of Clostridium innocuum in the human gut microbiota 94%
Similar papers in this journal
- Large-scale genomic analysis reveals the distribution and diversity of Type VI Secretion Systems in Escherichia coli 95%
- A comparison of short- and long-read whole genome sequencing for microbial pathogen epidemiology 94%
- Genomic characterization of the C. tuberculostearicum species complex, a ubiquitous member of the human skin microbiome 94%
Similar papers in this journal
- Species-wide phylogenomics of the Staphylococcus aureus agr operon reveals convergent evolution of frameshift mutations 96%
- Pangenome evaluation of gene essentiality in Streptococcus pyogenes 94%
- Establishment of a publicly available core genome multilocus sequence typing scheme for Clostridium perfringens 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.