Microbial Genomics
● Microbiology Society
All preprints, ranked by how well they match Microbial Genomics's content profile, based on 225 papers previously published here. The average preprint has a 0.16% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Hadjirin, N. F.; Yassine, I.; Bray, J. E.; Maiden, M. C. J.; Jolley, K. A.; Brueggemann, A. B.
Show abstract
Staphylococcus aureus infects both humans and animals, and antimicrobial resistance, including multidrug resistance, complicates treatment of S. aureus infections. Understanding S. aureus population structure and the distribution of genetic lineages is central to understanding the biology, epidemiology, and pathogenesis of this organism. This study exploited a large, publicly available dataset of nearly 27,000 S. aureus genomes to: i) develop a core genome multilocus sequence type (cgMLST) scheme; ii) stratify hierarchical clusters based on allelic similarity thresholds; and iii) define the clusters with a life identification number (LIN) code classification system. The cgMLST scheme characterised allelic variation at 1,716 core gene loci, and 13 classification thresholds were defined, which discriminated S. aureus variants across a range of genetic similarity thresholds. LIN code lineages and clonal complexes defined by seven-locus multilocus sequence typing were highly concordant, but the LIN codes permitted a wider range of genetic discrimination among S. aureus genomes. This S. aureus cgMLST scheme and LIN code system is a high-resolution, stable genotyping tool that enables detailed genomic analyses of S. aureus.
Rajendran, S.; Nagarajan, S.; MOHAN S., S.
Show abstract
Background: Recurrent urinary tract infections(rUTI) represent a major clinical challenge due to persistent clinical symptoms, repeated antibiotic exposure, and increased risk of multidrug resistance. Further clinical management of rUTI remains challenging, as existing diagnostic and treatment guidelines are largely designed for uncomplicated, acute infections. Though uropathogenic Escherichia coli (UPEC) is the predominant cause of community-acquired UTIs, pathogen-derived genomic features that may predispose certain E. coli strains to repeatedly establish infection are not fully understood. Methods: To comprehensively dissect distinct genetic signals across genomic compartments that distinguish rUTI-associated isolates from those causing sporadic infection, the pan-genome analysis in three different frameworks (i) Combined genomes (chromosome + plasmid), (ii) bacterial chromosomes only and (iii) plasmid-only was conducted. A comprehensive evaluation of population structure was performed using Gubbins, recombination-aware phylogeny IQTree, phylogroup distribution, pan-genome openness using Heaps law, and plasmidome architecture using MOBSUITE. Findings: Supervised machine learning models showed that the highest discriminatory performance was achieved using the combined genomic dataset (accuracy ~0.98), and integration of feature-selected genes with PanGWAS (Pyseer and Scoary) identified a robust set of recurrence-associated genes, namely cbtA, cbeA, and ldrD, which were consistently detected across machine learning and association frameworks. Subsequent association rule mining further revealed cooperative gene networks enriched in rUTI isolates, particularly involving toxin-antitoxin modules and metabolic regulators. Interpretation: Overall, this integrated ML-PanGWAS approach demonstrates that rUTI is a lineage-independent, polygenic phenotype encoded within a combined chromosomal-plasmid genomic context, providing new insights into the bacterial genomic architecture underlying recurrent disease and offering candidate biomarkers for future diagnostic and therapeutic development.
Parfitt, K. M.; Pascoe, B.; Jolley, K. A.; Douglas, A.; Goforth, M. P.; Sheppard, S. K.; Maiden, M. C. J.; Colles, F. M.
Show abstract
Campylobacter remains the leading cause of bacterial gastroenteritis worldwide, with C. jejuni accounting for around 90% of infection and C. coli accounting for most of the rest. Seven-locus multilocus sequence typing (MLST) has improved our understanding of host association and population structure, whilst core genome MLST (cgMLST), enables investigation of transmission events at high-resolution. However, the lack of a stable and standardised nomenclature for clustering of cgMLST data has limited reproducibility and long-term comparability between studies. Here we introduce a joint, hierarchical Life Identification Number (LIN) code system that provides reproducible, multi-level genomic identifiers for C. jejuni and C. coli lineages. Using an updated cgMLST v2 scheme (1,142 loci) and globally representative datasets of high-quality genomes selected from over 53,000 assemblies in the Campylobacter PubMLST database (https://pubmlst.org/organisms/campylobacter-jejunicoli), we firstly defined LIN codes on a dataset of 5,664 genomes. Pairwise allelic distances were computed using MSTclust, and 18 nested thresholds were defined through silhouette, adjusted Wallace and adjusted Rand Index (ARI) statistics to capture the population structure from species to outbreak level resolution. The LIN thresholds were then validated using a second dataset of 1,781 genomes from PubMLST and applied to a large water-associated outbreak dataset from New Zealand in 2016, containing clinical and ecological genomes. Further application of LIN codes was demonstrated by analyses of the C. jejuni ST-21 clonal complex and ST-6175 isolates, as well as the broader population structure of C. coli, using data from PubMLST. Across all datasets, LIN clusters were stable, largely monophyletic, and back-compatible with existing nomenclature, accurately distinguishing host-adapted and outbreak-associated lineages. By embedding cgMLST data within a stable and scalable nomenclature, the Campylobacter LIN system delivers consistent, automated genome-to-lineage assignment. This unified framework bridges population genetics and applied surveillance, enabling robust, real-time comparison of Campylobacter isolates across sources, studies, and time. Impact statementHuman cases of Campylobacter worldwide continue unabated. Tracing the source of Campylobacter infection is particularly challenging given the sporadic or multi-source nature of outbreaks, with potential transmission from foodborne, animal or environmental sources. Seven-locus MLST has greatly improved our broad understanding of Campylobacter population structure. However, whilst high-resolution cgMLST alleles and STs themselves do not change, longitudinal cluster analyses of cgMLST data have lacked a stable nomenclature, rendering them unsuitable for robust and comparable surveillance over time. Life Identification Number (LIN) codes provide a solution to this problem, establishing an automated and scalable nomenclature derived directly from cgMLST profiles, that is stable over time. We have implemented a joint C. jejuni and C. coli LIN code scheme in PubMLST, with scripts for real-time lineage assignment. LIN codes are back-compatible with existing MLST nomenclature, and we demonstrate their added practical value for exploring population structure and high-resolution outbreak investigation. LIN codes support surveillance of Campylobacter in a One Health context, by enabling consistent typing at multiple levels across different sources, laboratories and time. Data summary1. The isolate collections used to develop the LIN codes are publicly available and searchable as individual projects on the PubMLST database (https://pubmlst.org). O_LILIN code development (Dataset 1) (n=5,664 isolates, up to 200 isolates per clonal complex) C_LIO_LILIN code validation (Dataset 2) (n=1,781 isolates), up to 50 isolates per clonal complex C_LIO_LIOutbreak investigation (Dataset 3): New Zealand 2016 Havelock North waterborne outbreak, Gilpin et al (n=161 isolates) [1] C_LIO_LIPopulation structure exploration (clonal complex) (Dataset 4): ST-21 complex (n=1800 isolates, up to 100 isolates randomly selected from each country) C_LIO_LIPopulation structure exploration (sequence type) (Dataset 5): ST-6175 (n=321 isolates, genomes with good cgMLST v2 annotation) C_LI 2. The software for LIN code development is publicly available as follows: O_LIMSTclust for pairwise distance matrices https://gitlab.pasteur.fr/GIPhy/MSTclust [2] C_LIO_LIPython script to define LIN codes in a local dataset; (https://gitlab.pasteur.fr/BEBP/LINcoding) C_LIO_LIBIGSdb Perl script to define LIN codes from cgMLST profiles on the PubMLST database; (https://github.com/kjolley/BIGSdb/blob/develop/scripts/maintenance/lincodes.pl) C_LI
Sim, E. M.; Fong, W.; Suster, C.; Agius, J. E.; Chandra, S.; Suliman, B.; Wang, Q.; Ngo, C.; Finemore, C.; Chen, S. C.-A.; Basile, K.; Sintchenko, V.
Show abstract
Shiga Toxin (Stx) producing Escherichia coli (STEC) is a subset of pathogenic E. coli that can produce two types of Stx, Stx1 and Stx2, which can be further subtyped into four and 15 subtypes respectively. Not all subtypes, however, are equal in virulence potential, and the risk of severe disease including haemolytic uraemic syndrome has been linked to certain Stx2 subtypes e.g. Stx2a, Stx2d, highlighting the importance to survey stx subtypes. Previously, we developed a STEC virulence barcode to capture pertinent information on virulence genes to infer pathogenic potential. However, the process required multiple manual curation steps to determine the barcode. Here we introduce STECode, a bioinformatic tool to automate the STEC virulence barcode generation from sequencing reads or genomic assemblies. The development, and validation of STECode is described using a set of publicly available completed STEC genomes, along with their corresponding short reads. STECode was applied to interrogate the virulence landscape and molecular epidemiology of human STEC isolated during the period of the international border closures related to COVID-19 in the state of New South Wales, Australia. Impact statementWhole genome sequencing has been used to great effect in the genomic surveillance of STEC for public health purposes via the tracking of outbreaks. With STECode, we present a method to generate a STEC virulence barcode which captures pertinent subtyping information, useful for genomic inference of pathogenic potential. A key blind spot generated in short-read sequencing is the inability to detect the presence of multiple, isogenic stx copies in STEC. STECode mitigates this by inferring and reporting on the possibility of this occurrence. We envisage that this tool will value-add current genomic surveillance workflows through the ability to infer pathogenic potential.
Nagy, D.; Pennetta, V.; Rodger, G.; Hopkins, K.; Jones, C. R.; The NEKSUS Consortium, ; Hopkins, S.; Crook, D.; Walker, A. S.; Robotham, J.; Hopkins, K. L.; Ledda, A.; Williams, D.; Hope, R.; Brown, C. S.; Stoesser, N.; Lipworth, S.
Show abstract
Whole bacterial genome sequence reconstruction using Oxford Nanopore Technologies ("Nanopore") long-read only sequencing may offer a lower-cost, higher-throughput alternative for pathogen surveillance to hybrid assembly with recent improvements in Nanopore sequencing accuracy. We evaluated the accuracy, including plasmid reconstruction, of Nanopore long-read only genome assemblies of Enterobacterales. We sequenced 92 genomes from clinical Enterobacterales isolates, collected in England under a national surveillance program, with long-read Nanopore (R10.4.1, Dorado v5.0.0 super-high-accuracy basecalled) and short-read Illumina (NovaSeq) sequencing approaches. Genomes were assembled using three long-read only (Flye; Hybracter long; Autocycler), and three hybrid assemblers (Hybracter hybrid; Unicycler normal; bold). Three polishing modalities (Medaka v2 with subsampled or un-subsampled long-reads; Polypolish + Pypolca with short-reads) were investigated. Autocycler circularised the most chromosomes (87/92 [95%]). Plasmid sequence reconstruction was comparable between all assemblers except Flye, all recovering 90-96% of plasmids, although the ground truth was uncertain. Flye performed worse than other assemblers on almost all metrics. Autocycler + Medaka (un-subsampled long-reads) was the most accurate long-read only assembler/polisher combination, comparable to hybrid assemblies (median 0 [IQR:0-0] SNPs and 0 [IQR:0-1] indels per genome; quality value/Q score, 100 [IQR: 64-100]), with only 4/92 genome sequences having >10 SNPs/indels. Medaka polishing with un-subsampled long-reads resulted in small improvements in indels but not SNPs for both Flye and Autocycler assemblies. Seven-locus MLST, antimicrobial resistance, virulence, and stress gene annotation was equivalent across assembler/polisher combinations. Nanopore long-read only bacterial genome assembly with Autocycler combined with Medaka polishing (using un-subsampled reads) is similarly accurate and possibly more complete than hybrid assemblies, representing a viable alternative for incorporating high-quality genomic data, including plasmids, into Enterobacterales surveillance. Data SummaryNanopore long-reads and Illumina short-reads from the 92 Enterobacterales isolates from this study have been uploaded to ENA (BioProject accession: PRJEB93885). Code for the Nextflow assembly pipeline, downstream analysis scripts, and R statistical analysis scripts are available on GitHub (https://github.com/oxfordmmm/NEKSUS_ont_hybrid_assembly_comparison). The following supplementary data tables are available on FigShare (https://figshare.com/account/home#/projects/253775): O_LIENA Sample accessions and sample metadata (accessions_and_metadata.csv) C_LIO_LISeqkit stats summaries of the Illumina and Nanopore reads (raw_qc_sup.cav) C_LIO_LISummary of assembly contig features (contigs_summary_sup_cleaned.csv) C_LIO_LIPairwise mash distances between contigs (mash_cleaned.csv) C_LIO_LIPlasmids matching across different assemblers compared to the Hybracter (hybrid) and manually-curated reference sets (plasmids_match_hybracter_mash.csv; plasmids_match_manual_mash.csv, respectively) C_LIO_LISeven-locus multi-locus sequence type annotation (mlst_cleaned.csv) C_LIO_LICheckM2 summaries of assemblies (checkm2_cleaned.csv) C_LIO_LINucleotide-level accuracy of assemblies (SNP, Indels, and Quality value compared to short-read mapping; assembly_nucleotide_accuracy_cleaned.csv) C_LIO_LIBakta annotation (bakta_by_contig_cleaned.csv) C_LIO_LIAMRFinderPlus annotations of contigs (amrfinder_plus_cleaned.csv) C_LIO_LIMOB-suite annotation summaries of contigs (mobsuite_cleaned.csv) C_LI Impact StatementNanopore long-reads have historically been too error-prone to use alone for accurate bacterial genome assembly, necessitating additional Illumina short-reads to achieve structurally complete and accurate hybrid genome assemblies for public health surveillance. This increases cost and complexity. Previous studies have shown that recent improvements in Nanopore chemistry (R10.4.1 flowcell) and basecalling (super-high accuracy) allow high-quality long-read only assemblies on a small number of laboratory reference strains. This is the first evaluation, to our knowledge, to assess Nanopore long-read only genome assembly compared with hybrid assembly on a large number of clinical isolates. In addition, this is the first large-scale evaluation of the recently released automated consensus long-read assembly tool, Autocycler. We show that Autocycler long-read only assemblies are more structurally complete for chromosomal sequences, while reconstructing a similar number of plasmids to other long-read and hybrid assemblers. Most long-read polished, Autocycler-assembled genome sequences have 0 errors (median: 0 SNPs/indels) relative to a short-read polished (hybrid) Autocycler assemblies, enabling accurate annotation of key genes.
Bennett, R. J.; De Silva, P. M.; Bengtsson, R. J.; Horsburgh, M. J.; Blower, T. R.; Baker, K. S.
Show abstract
Bacteria of the genus Shigella are a major contributor to the global diarrhoeal disease burden causing >200,000 deaths per annum globally where S. flexneri is the major pathogenic species. Increasing antimicrobial resistance (AMR) in Shigella and the lack of a licenced vaccine has led WHO to recognise Shigella as a priority organism for the development of new antimicrobials. Understanding what drives the long-term persistence and success of this pathogen is critical for ongoing shigellosis management and is relevant for other enteric bacteria. To identify key genetic drivers of Shigella evolution over the past 100 years, we analysed S. flexneri from the historical Murray collection (n=45, isolated between 1917-1954) alongside a comparatively modern collection (n=262, isolated between 1950-2011) using a novel approach called temporal genome-wide association study (tGWAS). We identified SNPs (n=94), COGs (n=359) and significant kmers within 48 genes significantly positively associated with time. These included T3SS encoding genes, proteins involved in intracellular competition, acquired antimicrobial resistance genes, insertion sequences, and genes of unknown function (28%, 49/172 of those hits investigated). Among the unknown proteins we identified a novel plasmid borne putative adhesin, named Stv. Genomic epidemiological analyses reveal that Stv was associated with clonal expansions of multiple phylogroups of S. flexneri and its acquisition predates multidrug resistance acquisition and the global dissemination of Lineage III S. sonnei. Stv, and close relatives, are widely distributed in other Enterobactericeae and bacteria, indicating that its importance likely extends beyond shigellae. This work highlights the effectiveness of using tGWAS on historical isolate collections for identifying novel contributors to pathogen success over time. This approach is readily translatable to other pathogens and our application in Shigella identified Stv, a putative adhesin and potential drug target that is widely distributed across the AMR priority group Enterobacteriaceae. Author SummaryShigellosis is a leading cause of diarrhoeal disease worldwide and is represented among the multiple Enterobacteriaceae which WHO have declared as priority pathogens for which new antimicrobials are urgently needed. The majority of shigellosis is caused by the species Shigella flexneri and Shigella sonnei. In this study we collated S. flexneri isolates that spanned a 94-year period, encompassing the pre- and post-antibiotic era to implement a novel bioinformatic technique, temporal GWAS (tGWAS), to identify key factors of pathogen success during this time period. Alongside recovering AMR and virulence genes, we also identified a novel, mobilisable putative adhesin, named Stv herein, which appeared to contribute to clonal expansions across multiple Shigella species and is present across the broader Enterobacteriaceae. Our results indicate the potential importance of Stv in controlling Shigella and other infections, and the validity of a tGWAS approach for identifying biological drivers underpinning the evolution and expansion of AMR pathogens over time.
Slow, S.; Anderson, T.; Murdoch, D.; Bloomfield, S.; Winter, D.; Biggs, P.
Show abstract
Legionella longbeachae is an environmental bacterium that is commonly found in soil and composted plant material. In New Zealand (NZ) it is the most clinically significant Legionella species causing around two-thirds of all notified cases of Legionnaires disease. Here we report the sequencing and analysis of the geo-temporal genetic diversity of 54 L. longbeachae serogroup 1 (sg1) clinical isolates that were derived from cases from around NZ over a 22-year period, including one complete genome and its associated methylome. Our complete genome consisted of a 4.1 Mb chromosome and a 108 kb plasmid. The genome was highly methylated with two known epigenetic modifications, m4C and m6A, occurring in particular sequence motifs within the genome. Phylogenetic analysis demonstrated the 54 sg1 isolates belonged to two main clades that last shared a common ancestor between 108 BCE and 1608 CE. These isolates also showed diversity at the genome-structural level, with large-scale arrangements occurring in some regions of the chromosome and evidence of extensive chromosomal and plasmid recombination. This includes the presence of plasmids derived from recombination and horizontal gene transfer between various Legionella species, indicating there has been both intra-species and inter-species gene flow. However, because similar plasmids were found among isolates within each clade, plasmid recombination events may pre-empt the emergence of new L. longbeachae strains. Our high-quality reference genome and extensive genetic diversity data will serve as a platform for future work linking genetic, epigenetic and functional diversity in this globally important emerging environmental pathogen. Author SummaryLegionnaires disease is a serious, sometimes fatal pneumonia caused by bacteria of the genus Legionella. In New Zealand, the species that causes the majority of disease is Legionella longbeachae. Although the analyses of pathogenic bacterial genomes is an important tool for unravelling evolutionary relationships and identifying genes and pathways that are associated with their disease-causing ability, until recently genomic data for L. longbeachae has been sparse. Here, we conducted a large-scale genomic analysis of 54 L. longbeachae isolates that had been obtained from people hospitalised with Legionnaires disease between 1993 and 2015 from 8 regions around New Zealand. Based on our genome analysis the isolates could be divided into two main groups that persisted over time and last shared a common ancestor up to 1700 years ago. Analysis of the bacterial chromosome revealed areas of high modification through the addition of methyl groups and these were associated with particular DNA sequence motifs. We also found there have been large-scale rearrangements in some regions of the chromosome, producing variability between the different L. longbeacahe strains, as well as evidence of gene-flow between the various Legionella species via the exchange of plasmid DNA.
Lam, M. M. C.; Wick, R. R.; Holt, K. E.; Wyres, K. L.
Show abstract
The outer polysaccharide capsule and lipopolysaccharide antigens are key targets for novel control strategies targeting Klebsiella pneumoniae and related taxa from the K. pneumoniae species complex (KpSC), including vaccines, phage and monoclonal antibody therapies. Given the importance and growing interest in these highly diverse surface antigens, we had previously developed Kaptive, a tool for rapidly identifying and typing capsule (K) and outer lipopolysaccharide (O) loci from whole genome sequence data. Here, we report two significant updates, now freely available in Kaptive 2.0 (github.com/katholt/kaptive); i) the addition of 16 novel K locus sequences to the K locus reference database following an extensive search of >17,000 KpSC genomes; and ii) enhanced O locus typing to enable prediction of the clinically relevant O2 antigen (sub)types, for which the genetic determinants have been recently described. We applied Kaptive 2.0 to a curated dataset of >12,000 public KpSC genomes to explore for the first time the distribution of predicted O (sub)types across species, sampling niches and clones, which highlighted key differences in the distributions that warrant further investigation. As the uptake of genomic surveillance approaches continues to expand globally, the application of Kaptive 2.0 will generate novel insights essential for the design of effective KpSC control strategies. Significance as a BioResource to the communityKlebsiella pneumoniae is a major cause of bacterial healthcare associated infections globally, with increasing rates of antimicrobial resistance, including strains with resistance to the drugs of last resort. The latter have therefore been flagged as priority pathogens for the development of novel control strategies. K. pneumoniae produce two key surface antigen sugars (capsular polysaccharide and lipopolysaccharide (LPS)) that are immunogenic and targets for novel controls such as a vaccines and phage therapy. However, there is substantial antigenic diversity in the population and relatively little is understood about the distribution of antigen types geographically and among strains causing different types of infections. Whereas laboratory-based antigen typing is difficult and rarely performed, information about the relevant synthesis loci can be readily extracted from whole genome sequence data. We have previously developed Kaptive, a freely available tool for rapid typing of Klebsiella capsule and LPS loci from genome sequences. Kaptive is now used widely in the global research community and has facilitated new insights into Klebsiella capsule and LPS diversity. Here we present an update to Kaptive facilitating i) the identification of 16 additional novel capsule loci, and ii) the prediction of immunologically relevant LPS O2 antigen subtypes. These updates will enable enhanced sero-epidemiological surveillance for K. pneumoniae, to inform the design of vaccines and other novel Klebsiella control strategies. Data summaryO_LIThe updated code and reference databases for Kaptive are available at https://github.com/katholt/Kaptive C_LIO_LIGenome accessions from which reference sequences of novel K loci were defined are listed in Supplementary Table 1, and genomes from which these loci were detected (along with the corresponding Kaptive output) are listed in Supplementary Table 2. C_LIO_LIAccessions for the genomes screened for O types/subtypes (along with the corresponding Kaptive output) are listed in Supplementary Table 3. C_LI The authors confirm all supporting data, code and protocols have been provided within the article or through supplementary data files. Repositories1.1 RepositoriesGenome sequence from which the novel K locus KL182 was defined has been deposited under the accession JAJHNT000000000.
Nair, S.; Barker, C. R.; Bird, M.; Ledda, A.; Collins, C.; Morrison, R.; Greig, D. R.; Rodwell, E. V.; Painset, A.; Crewdson, A.; Jenkins, C.; Chattaway, M. A.; Didelot, X.; Ribeca, P.
Show abstract
The life cycle of certain bacteriophages involves their maintenance within the bacterial cell as extrachromosomal elements, complete with replication and partitioning systems. These "phage-plasmids" (P-Ps) are distributed widely among bacterial phyla but are not typically included during genomic surveillance studies, and previous reports do not consider the context of their host strain diversity. We recently identified a P1-like P-P carrying the blaCTX-M-15 resistance gene in Salmonella enterica serovar Typhi, prompting a subsequent investigation into the overall frequency of P-Ps among gastrointestinal bacteria under routine genomic surveillance within England. We expanded our study to include P-P groups known to be associated with Enterobacteriaceae (P1, D6, SSU5, N15) and scanned a collection of 66,856 genomes of diarrhoeagenic Escherichia spp., Shigella spp. and S. enterica. All four P-P groups were detected in our dataset, totalling 9% of E. coli and Shigella genomes and 2% of S. enterica genomes. A small subset harboured two distinct P-P groups. P1-like P-Ps, some of which carried a cytotoxic necrotising factor, were predominantly associated with Shiga toxin-producing E. coli strains of public health concern, including clonal complexes CC11 (O157:H7), CC29 (O26:H11) and CC165. In contrast, SSU5 group P-Ps were linked with multidrug-resistant lineages of both S. Typhi and Shigella sonnei. We found multiple antimicrobial resistance genes inserted into P-Ps via transposons, integrons and insertion sequences, and numerous defence/anti-defence systems. There was evidence of vertical transmission but also many links across species, time and geography, indicating that horizontal transfer is occurring regularly. Phage-plasmids are often described only as cryptic elements or not detected during genomic surveillance. We show that P-Ps are associated with clinically relevant lineages of human pathogens and can acquire accessory genes that may impact on disease severity and therefore should play a more prominent role in pathogen surveillance and epidemiology.
Ashcroft, M.; McGarry, N.; Stanton, T. D.; Hoyles, L.; Holt, K. E.; Wyres, K. L.
Show abstract
The Klebsiella oxytoca Species Complex (SC) represents an emerging healthcare-associated group of opportunistic pathogens. The capsular polysaccharide is a virulence determinant and target for novel vaccines, monoclonal antibodies and phage therapy. In the absence of broadly accessible phenotyping techniques, prediction of capsule types from whole-genome sequence data is critical for understanding capsule diversity and epidemiology, and to prioritise capsule types as targets for novel anti-K. oxytoca SC interventions. Here we present the first comprehensive capsule synthesis locus (K locus) database targeted for the K. oxytoca SC, comprising 88 distinct loci defined by gene content and which is compatible with the rapid genome typing tool Kaptive. The database provides high coverage of publicly available K. oxytoca SC genomes (97.6% of 2,244 genomes, dereplicated from a total of 4,055), and the typing rate is significantly higher than that achieved with the pre-existing Klebsiella K locus database (97.6% vs 50.3%, p <0.0001), which primarily targets the Klebsiella pneumoniae SC. We demonstrate the utility of the novel K. oxytoca SC database by application to three diverse clinical K. oxytoca SC isolate collections (n=61 to 102 genomes each), suggesting a high diversity of K types. The novel K. oxytoca SC K locus database (github.com/klebgenomics/KoSC-surface-antigen-loci) will provide a key resource to support larger systematic studies and ongoing genomics surveillance efforts for the K. oxytoca SC. IMPACT STATEMENTMembers of the Klebsiella oxytoca Species Complex (SC) are an emerging cause of infections in humans and are frequently associated with antimicrobial resistance. Klebsiella species produce two key surface antigen sugars (capsular polysaccharide and lipopolysaccharide) that are immunogenic and are targets for novel control strategies such as vaccines and phage therapy. Phenotypic typing of these surface antigen sugars (serotyping) is costly and laborious, with genotyping (predicting the serotype from whole-genome sequence data) a useful alternative. Here, we present a curated capsule (K) locus reference database for the K. oxytoca SC, which represents a useful tool to assist in the epidemiological surveillance of this emerging pathogen. DATA SUMMARYAll Klebsiella oxytoca Species Complex genomes used in this work were publicly available, with accession details listed in Supplementary Tables 1, 2 and 3. The K. oxytoca Species Complex K locus reference database is available under GNU public license at github.com/klebgenomics/KoSC-surface-antigen-loci.
Rodrigues, C.; Lanza, V. F.; Peixe, L. V.; Coque, T. M.; Novais, A.
Show abstract
The increasing worldwide spread of multidrug-resistant (MDR) Kp is largely driven by high-risk sublineages, some of them well-characterised such as Clonal Group (CG) 258, CG147 or CG307. MDR Kp Sequence-Type (ST) 14 and ST15 have been described worldwide causing frequent outbreaks of CTX-M-15 and/or carbapenemase producers. However, their phylogeny, population structure and global dynamics remain unclear. Here, we clarify the phylogenetic structure and evolvability of CG14 and CG15 Kp by analysing the CG14 and CG15 genomes available in public databases (n=481, November 2019) and de novo sequences representing main sublineages circulating in Portugal (n=9). Deduplicated genomes (n=235) were used to infer temporal phylogenetic evolution and to compare their capsular locus (KL), resistome, virulome and plasmidome using high-resolution tools. Phylogenetic analysis supported independent evolution of CG14 and CG15 within two distinct clades and 4 main subclades which are mainly defined according to the KL and the accessory genome. Within CG14, two large monophyletic subclades, KL16 (14%) and KL2 (86%), presumptively emerged around 1937 and 1942, respectively. Sixty-five percent of CG14 carried genes encoding ESBL, AmpC and/or carbapenemases and, remarkably, they were mainly observed in the KL2 subclade. The CG15 clade was segregated in two major subclades. One was represented by KL24 (42%) and KL112 (36%), the latter one diverging from KL24 around 1981, and the other comprised KL19 and other KL-types (16%). Of note, most CG15 genomes contained genes encoding ESBL, AmpC and/or carbapenemases (n=148, 87%) and displayed a characteristic set of mutations in regions encoding quinolone resistance (QRDR, GyrA83F/GyrA87A/ParC80I). Plasmidome analysis revealed 2463 plasmids grouped in 27 predominant plasmid groups (PG) with a high degree of recombination, including particularly pervasive F-type (n=10) and Col (n=10) plasmids. Whereas blaCTX-M-15 was linked to a high diversity of mosaic plasmids, other ARGs were confined to particular plasmids (e.g. blaOXA-48-IncL; blaCMY/TEM-24-IncC). This study firstly demonstrates an independent evolutionary trajectory for CG15 and CG14, and suggests how the acquisition of specific KL, QRDR mutations (CG15) and ARGs in highly recombinant plasmids could have shaped the expansion and diversification of particular subclades (CG14-KL2, CG15-KL24/KL112). IMPORTANCEKlebsiella pneumoniae (Kp) represents a major threat in the burden of antimicrobial resistance (AMR). Phylogenetic approaches to explain the phylogeny, emergence and evolution of certain multidrug resistant populations have mainly focused on core-genome approaches while variation in the accessory genome and the plasmidome have been long overlooked. In this study, we provide unique insights into the phylogenetic evolution and plasmidome of two intriguing and yet uncharacterized clonal groups (CGs), the CG14 and CG15, which have contributed to the global dissemination of contemporaneous {beta}-lactamases. Our results point-out an independent evolution of these two CGs and highlight the existence of different clades structured by the capsular-type and the accessory genome. Moreover, the contribution of a turbulent flux of plasmids (especially multireplicon F type and Col) and adaptive traits (antibiotic resistance and metal tolerance genes) to the pangenome, reflect the exposure and adaptation of Kp under different selective pressures.
Krisna, M. A.; Jolley, K. A.; Monteith, W.; Boubour, A.; Hamers, R. L.; Brueggemann, A. B.; Harrison, O. B.; Maiden, M. C. J.
Show abstract
2.Haemophilus influenzae is part of the human nasopharyngeal microbiota and a pathogen causing invasive disease. The extensive genetic diversity observed in H. influenzae necessitates discriminatory analytical approaches to evaluate its population structure. This study developed a core genome MLST (cgMLST) scheme for H. influenzae using pangenome analysis tools and validated the cgMLST scheme using datasets consisting of complete reference genomes (N=14) and high-quality draft H. influenzae genomes (N=2,297). The draft genome dataset was divided into a development (N=921) and a validation dataset (N=1,376). The development dataset was used to identify potential core genes with the validation dataset used to refine the final core gene list to ensure the reliability of the proposed cgMLST scheme. Functional classifications were made for all resulting core genes. Phylogenetic analyses were performed using both allelic profiles and nucleotide sequence alignments of the core genome to test congruence, as assessed by Spearmans correlation and Ordinary Least Square linear regression tests. Preliminary analyses using the development dataset identified 1,067 core genes, which were refined to 1,037 with the validation dataset. More than 70% of core genes were predicted to encode proteins essential for metabolism or genetic information processing. Phylogenetic and statistical analyses indicated that the core genome allelic profile accurately represented phylogenetic relatedness among the isolates (R2 = 0.945). We used this cgMLST scheme to define a high-resolution population structure for H. influenzae, which enhances the genomic analysis of this clinically relevant human pathogen. 3. Impact statementDiscriminating H. influenzae variants and evaluating population structure has been challenging and largely unstandardised. To address this, we have developed a cgMLST scheme for H. influenzae. Since an accurate typing approach relies on precise reflection of the underlying population structure, we explored various methods to define the scheme. The core genes included in this scheme were predicted to encode functions in essential biological pathways, such as metabolism and genetic information processing, and could be reliably assembled from short-read sequence data. Single-linkage clustering, based on core genome allelic profiles, showed high congruence to genealogy reconstructed by Maximum-Likelihood (ML) methods from the core genome nucleotide alignment. The cgMLST scheme v1 enables rapid and accurate depiction of high-resolution H. influenzae population structure, and making this scheme accessible via the PubMLST database, ensures that microbiology reference laboratories and public health authorities worldwide can use it for genomic surveillance. 4. Data summaryThe H. influenzae cgMLST scheme is accessible via https://pubmlst.org/organisms/haemophilus-influenzae. The list of isolate IDs available publicly from pubmlst.org is provided in Supplementary File 1. The pipeline for cgMLST scheme development and validation is published at https://www.protocols.io/private/EF6DB7FE429311EEB8630A58A9FEAC02. All in-house R and Python scripts for data processing and analysis are available from https://gitfront.io/r/user-4399403/ZHt8DArALHcY/cgmlst-hinf/.
Stoesser, N.; Phan, H. T.; Seale, A. C.; Aiken, Z.; Thomas, S.; Smith, M.; Wyllie, D.; George, R.; Sebra, R.; Mathers, A. J.; Vaughan, A.; Peto, T. E.; Ellington, M. J.; Hopkins, K. L.; Crook, D. W.; Orlek, A.; Welfare, W.; Cawthorne, J.; Lenney, C.; Dodgson, A.; Woodford, N.; Walker, A. S.; TRACE Investigators' Group,
Show abstract
Carbapenem resistance in Enterobacterales is a public health threat. Klebsiella pneumoniae carbapenemase (encoded by alleles of the blaKPC family) is one of the commonest transmissible carbapenem resistance mechanisms worldwide. The dissemination of blaKPC has historically been associated with distinct K. pneumoniae lineages (clonal group 258 [CG258]), a particular plasmid family (pKpQIL), and a composite transposon (Tn4401). In the UK, blaKPC has caused a large-scale, persistent outbreak focused on hospitals in North-West England. This outbreak has evolved to be polyclonal and poly-species, but the genetic mechanisms underpinning this evolution have not been elucidated in detail; this study used short-read whole genome sequencing of 604 blaKPC-positive isolates (Illumina) and long-read assembly (PacBio)/polishing (Illumina) of 21 isolates for characterisation. We observed the dissemination of blaKPC (predominantly blaKPC-2; 573/604 [95%] isolates) across eight species and more than 100 known sequence types. Although there was some variation at the transposon level (mostly Tn4401a, 584/604 (97%) isolates; predominantly with ATTGA-ATTGA target site duplications, 465/604 [77%] isolates), blaKPC spread appears to have been supported by highly fluid, modular exchange of larger genetic segments amongst plasmid populations dominated by IncFIB (580/604 isolates), IncFII (545/604 isolates) and IncR replicons (252/604 isolates). The subset of reconstructed plasmid sequences also highlighted modular exchange amongst non-blaKPC and blaKPC plasmids, and the common presence of multiple replicons within blaKPC plasmid structures (>60%). The substantial genomic plasticity observed has important implications for our understanding of the epidemiology of transmissible carbapenem resistance in Enterobacterales, for the implementation of adequate surveillance approaches, and for control.\n\nIMPORTANCEAntimicrobial resistance is a major threat to the management of infections, and resistance to carbapenems, one of the \"last line\" antibiotics available for managing drug-resistant infections, is a significant problem. This study used large-scale whole genome sequencing over a five-year period in the UK to highlight the complexity of genetic structures facilitating the spread of an important carbapenem resistance gene (blaKPC) amongst a number of bacterial species that cause disease in humans. In contrast to a recent pan-European study from 2012-2013(1), which demonstrated the major role of spread of clonal blaKPC-Klebsiella pneumoniae lineages in continental Europe, our study highlights the substantial plasticity in genetic mechanisms underpinning the dissemination of blaKPC. This genetic flux has important implications for: the surveillance of drug resistance (i.e. making surveillance more difficult); detection of outbreaks and tracking hospital transmission; generalizability of surveillance findings over time and for different regions; and for the implementation and evaluation of control interventions.
MacFadyen, A. C.
Show abstract
Staphylococcal cassette chromosome (SCC) elements are mobile genetic elements that integrate at the rlmH gene and are predominantly responsible for methicillin resistance in staphylococci. Although SCCmec typing tools exist, none can extract the element sequence itself or explicitly classify SCC elements that lack methicillin resistance genes. Here we present SCCmecExtractor, a lightweight Python toolkit that identifies SCC element boundaries through degenerate attachment site (att) pattern matching, extracts complete elements from whole-genome assemblies and characterises their mec and ccr gene content. Benchmarking on 7,297 genomes spanning 70 species across Staphylococcus and Mammaliicoccus demonstrated 100% typing concordance with the sccmec tool1 on 1,454 S. aureus genomes. The tool extracted 1,562 SCC elements, from 1,454 S. aureus, 5,295 non-aureus Staphylococcus and 548 Mammaliicoccus genomes, achieving effective extraction rates (excluding assembly-limited genomes and those lacking valid ccr pairs) of 87.3% for S. aureus, 58.8% for non-aureus Staphylococcus, and 61.9% for Mammaliicoccus. Notably, 616 of the 1,562 extracted elements (39.4%) were non-mec SCC elements lacking methicillin resistance genes, a class of mobile element often overlooked. Non-mec SCC prevalence increased from 12.2% in S. aureus to 55.6% in non-aureus Staphylococcus and 76.0% in Mammaliicoccus, revealing a substantial reservoir of SCC diversity beyond methicillin resistance. SCCmecExtractor is freely available via PyPI, Docker and Singularity under an MIT licence. Impact StatementStaphylococcal cassette chromosome (SCC) elements are mobile genetic elements responsible for methicillin resistance in staphylococci and are central to methicillin resistant Staphylococcus aureus (MRSA) epidemiology. Existing tools focus on typing SCCmec from assemblies but cannot extract the element itself, limiting our ability to comprehensively monitor and examine these elements. SCCmecExtractor is a lightweight, portable tool that detects the attachment sites, required by SCC elements to integrate into the genome, extracts the SCC element, both mec gene carrying and not, and characterises their gene content. Applied across 7,297 genomes spanning two genera, we demonstrate that non-mec SCC elements are the dominant SCC class outside S. aureus, a finding enabled by systematic extraction and classification of SCC elements regardless of mec gene content. SCCmecExtractor provides the research community with an accessible, confidence-first approach (based on biology) to SCC element analysis across all staphylococci and mammaliicocci species. Data SummaryThe code for this pipeline is available at: https://github.com/AlisonMacFadyen/SCCmecExtractor, with a Docker image available at: https://hub.docker.com/repository/docker/alisonmacfadyen/sccmecextractor and PyPi package at: https://pypi.org/project/sccmecextractor/. All reference databases are bundled with the tool. Benchmarking genome accessions: 1,454 S. aureus, 5,295 non-aureus Staphylococcus, and 548 Mammaliicoccus genomes from NCBI. A complete list of genome accessions is provided as supplementary data (Supplementary Table S1). Extracted SCC elements can be obtained from Zenodo: 10.5281/zenodo.19355206
Parker, M. J.; Hopkins, K. M. V.; Chau, K. K.; Cregan, J.; Oakley, S.; Barrett, L.; Jeffery, K.; Butcher, L.; Paulus, S.; Young, B. C.; Eyre, D. W.; Fowler, P. W.; Stoesser, N.; Sanderson, N. D.; Bejon, P.
Show abstract
Antimicrobial resistance genes (ARGs) can spread via horizontal transfer or clonal expansion. We investigated the genomic epidemiology of extended-spectrum beta-lactamase (ESBL)--producing Klebsiella pneumoniae and Escherichia coli in a neonatal unit. Between January and November 2023, 53 ESBL isolates were obtained from 23 neonates via routine screening and clinical sampling. Long-read nanopore sequencing identified blaCTX-M-15 as the dominant ESBL gene, alongside blaCTX-M-65 and blaCTX-M-27. Among 49 blaCTX-M-15 isolates, 33 carried the gene on plasmids and in the remainder it was located on the chromosome. ESBL isolates belonged to one of seven MLST sequence types; E. coli isolates were dominated by ST131/ST131-like/ST5640 lineages, while K. pneumoniae were predominantly ST13. Plasmids (ESBL and non-ESBL-associated) clustered into 18 communities, five of which contained plasmid-bearing blaCTX-M genes. The largest cluster comprised IncFIB(K) plasmids from K. pneumoniae ST13, although these were predicted as "non-mobilizable" by MOB-suite and belonged to isolates from a clonally disseminated strain. No evidence of blaCTX-M dissemination via shared plasmids was identified. Meanwhile, chromosomal phylogenetic analysis identified four distinct clonal clusters with [≤]7 SNP differences (n = 10, 5, 5, and 2 patients). In this setting, genomic analysis supported clonal dissemination of several blaCTX-M-associated strains as the main outbreak mechanism, affecting 20/23 neonates, rather than plasmid-mediated transmission.
David, S.; Mentasti, M.; Sands, K.; Portal, E.; Graham, L.; Watkins, J.; Williams, C.; Healy, B.; Spiller, O. B.; Aanensen, D.; Wootton, M. B.; Jones, L. S.
Show abstract
Rising rates of multi-drug resistant Klebsiella infections necessitate a comprehensive understanding of the major strains and plasmids driving spread of resistance elements. Here we analysed 540 Klebsiella isolates recovered from patients across Wales between 2007 and 2020 using combined short- and long-read sequencing approaches. We identified resistant clones that have spread within and between hospitals including the high-risk strain, sequence type (ST) 307, which acquired the blaOXA-244 carbapenemase gene on a pOXA-48-like plasmid. We found evidence that this strain, which caused an acute outbreak largely centred on a single hospital in 2019, had been circulating undetected across South Wales for several years prior to the outbreak. In addition to clonal transmission, our analyses revealed evidence for substantial plasmid spread, mostly notably involving blaKPC-2 and blaOXA-48-like (including blaOXA-244) carbapenemase genes that were found among many species and strain backgrounds. Two thirds (20/30) of the blaKPC-2 genes were carried on the Tn4401a transposon and associated with IncF plasmids. These were mostly recovered from patients in North Wales, reflecting an outward expansion of the plasmid-driven outbreak of blaKPC-2-producing Enterobacteriaceae in North-West England. 92.1% (105/114) of isolates with a blaOXA-48-like carbapenemase carried the gene on a pOXA-48-like plasmid. While this plasmid family is highly conserved, our analyses revealed novel accessory variation including integrations of additional resistance genes. We also identified multiple independent deletions involving the tra gene cluster among pOXA-48-like plasmids in the ST307 outbreak lineage. These resulted in loss of conjugative ability and signal adaptation of the plasmids to carriage by the host strain. Altogether, our study provides the first high resolution view of the diversity, transmission and evolutionary dynamics of major resistant clones and plasmids of Klebsiella in Wales and forms an important basis for ongoing surveillance efforts. Data SummaryAll raw short read sequence data and hybrid assemblies are available in the European Nucleotide Archive (ENA) under project accession PRJEB48990.
Zuza, A. M.; Pearse, O.; Domman, D. B.; Dyson, Z. A.; Kawaza, K.; Musicha, P.; Feasey, N. A.; Heinz, E.
Show abstract
BackgroundKlebsiella pneumoniae (Kpn) is an important cause of healthcare-associated infections (HAI). In low and middle-income countries, HAI due to Kpn disproportionally affects neonates. In this study, we investigated the genomic changes that occurred during long-term circulation of a Kpn ST39 clone, causing a disproportionate number of infections on the neonatal ward at a tertiary healthcare facility in Malawi in 2017. MethodsWe analyzed whole genome sequences of Klebsiella pneumoniae ST39 collected from Queen Elizabeth Central Hospital over a 20-year period, including generation of several high-quality hybrid genomes. We compared virulence markers, antibiotic resistance determinants, and mobile genetic elements, focusing on variable regions between strains from the outbreak clone in 2017 to genomes from other co-occurring ST39 lineages. ResultsWe identified eight variable genomic regions that demonstrate the plasticity of Kpn within-ST, including the role of bacteriophages in shaping the genome of ST39. ConclusionsThe analyzed Klebsiella pneumoniae ST39 lineages have a highly variable genome capable of incorporating large genomic regions during prolonged hospital circulation, which may offer a selective advantage in hospital environments and provide resistance to antimicrobial agents. Data summaryAll sequencing data is available in BioProject PRJEB102175; detailed accession numbers are provided in Table S1. The authors confirm all supporting data, code and protocols have been provided within the article or through supplementary data files.
Cummins, E.; Moran, R.; Snaith, A.; Hall, R.; Connor, C.; Dunn, S.; McNally, A.
Show abstract
The repeated emergence of multi-drug resistant (MDR) Escherichia coli clones is a threat to public health globally. In recent work, drug resistant E. coli were shown to be capable of displacing commensal E. coli in the human gut. Given the rapid colonisation observed in travel studies, it is possible that the presence of a type VI secretion system (T6SS) may be responsible for the rapid competitive advantage of drug resistant E. coli clones. We employed large scale genomic approaches to investigate this hypothesis. First, we searched for T6SS genes across a curated dataset of over 20,000 genomes representing the full phylogenetic diversity of E. coli. This revealed large, non-phylogenetic variation in the presence of T6SS genes. No association was found between T6SS gene carriage and MDR lineages. However, multiple clades containing MDR clones have lost essential structural T6SS genes. We characterised the T6SS loci of ST410 and ST131 and identified specific recombination and insertion events responsible for the parallel loss of essential T6SS genes in two MDR clones. Data SummaryThe genome sequence data generated in this study is publicly available from NCBI under BioProject PRJNA943186, alongside a complete assembly in GenBank under accessions CP120633-CP120634. All other sequence data used in this paper has been taken from ENA with the appropriate accession numbers listed within the methods section. The E. coli genome data sets used in this work are from a previous publication, the details of which can be found in the corresponding supplementary data files 10.6084/m9.figshare.21360108 [1]. Impact StatementEscherichia coli is a globally significant pathogen that causes the majority of urinary tract infections. Treatment of these infections is exacerbated by increasing levels of drug resistance. Pandemic multi-drug resistant (MDR) clones, such as ST131-C2/H30Rx, contribute significantly to global disease burden. MDR E. coli clones are able to colonise the human gut and displace the resident commensal E. coli. It is important to understand how this process occurs to better understand why these pathogens are so successful. Type VI secretion systems may be one of the antagonistic systems employed by E. coli in this process. Our findings provide the first detailed characterisation of the T6SS loci in ST410 and ST131 and shed light on events in the evolutionary pathways of the prominent MDR pathogens ST410-B4/H42RxC and ST131-C2/H30Rx.
O'Ferrall, A. M.; Lally, D.; Makaula, P.; Namacha, G.; Lewis, J. M.; Musicha, P.; Goodman, R. N.; Allman, E.; Moyo, S.; Waddington, C. S.; Kayuni, S. A.; Feasey, N. A.; Musaya, J.; Stothard, J. R.; Roberts, A. P.
Show abstract
Infection with extended-spectrum beta-lactamase-producing Escherichia coli (ESBL-Ec) is a global health concern that disproportionately affects sub-Saharan Africa (SSA). Gut mucosal colonisation is thought to precede invasive infection. Understanding ESBL-Ec colonisation and transmission across communities is therefore essential. We investigated the genomic epidemiology and spatial structure of 159 gut-colonising ESBL-Ec isolates from the faeces of 211 people in two rural Malawian villages using longitudinal sampling (2023-24), whole-genome sequencing and household mapping. Colonisation prevalence rose from 34.1% (95% CI: 27.8-41.0) to 54.2% (95% CI: 46.0-62.3) over one year. Isolates belonged to 33 sequence types (STs), most commonly ST38 and ST131, harbouring 46 distinct antimicrobial resistance gene types. Fifteen strains were identified in [≥]3 households that were typically separated by short geographic distances (<400 m). Of 190 pairwise comparisons between same-strain isolates from different households sampled concurrently within villages, 88.9% differed by [≤]10 single nucleotide polymorphisms, consistent with multi-household involvement in community transmission networks. Lineage-specific ST38 and ST131 network analyses linked rural isolates to urban Malawian isolates collected within the last decade. Our findings provide a transferable framework for inferring ESBL-Ec flow in community settings and highlight the need for One Health surveillance and improved sanitation infrastructure to limit transmission.
Vezina, B.; White, C.; Cooper, H.; Holt, K.; Hawkey, J.; Wyres, K. L.; Lam, M. M.
Show abstract
IntroductionKlebsiella pneumoniae is an opportunistic pathogen which causes a wide spectrum of infections within healthcare settings and the community. Four K. pneumoniae sub-lineages defined with cgMLST/LINcodes are known to cause distinct infections of the nasal and/or upper respiratory passages: SL91 and SL10031 (also referred to as subspecies ozaenae), SL10032 (subspecies rhinoscleromatis) and SL82. These sub-lineages have also demonstrated reduced carbon source utilisation, which in other species has been linked with high loads of insertion sequences (IS). MethodsWe performed comparative genomics, analysed IS and constructed genome-scale metabolic models for available public sequences from these four sub-lineages and compared them to other sub-lineages from the wider K. pneumoniae population. ResultsThe four focal sub-lineages displayed significantly higher IS loads (median range 88 to 120 per genome) compared to other K. pneumoniae sub-lineages (median range 12 to 73). Notably, each K. pneumoniae sub-lineage had unique IS profiles, consistent with distinct evolutionary trajectories of IS acquisition and expansion. Across sub-lineages, higher IS loads were inversely associated with the number of metabolic model genes per genome (R2 = 0.16, p <0.001), as well as predicted aerobic substrate utilisation for phosphorus sources (R2 = 0.39, p<0.001) as per a second-degree polynomial regression model (n = 1,664 genomes). Additionally, the four IS-dense sub-lineages displayed a combination of convergent sub-lineage-specific substrate utilisation losses including parallel loss of 3-Phospho-D-glycerate, D-Glycerate-2-phosphate, Phosphoenolpyruvate utilisation as carbon/phosphorus sources. Finally, inspection of IS insertion sites demonstrated frequent and non-destructive insertion next to transcriptional, carbohydrate and amino acid metabolism genes. ConclusionsIS loads were significantly and inversely associated with metabolic substrate usage within K. pneumoniae, whereby sub-lineages that had higher numbers of IS also had reduced metabolic capacity. We hypothesise that an insertional tolerance model explains these findings, whereby IS can only insert into "metabolically-tolerable" sites for the individual cell and any impacts on metabolism are not detrimental for survival.