Discordance between different bioinformatic methods for identifying resistance genes from short-read genomic data, with a focus on Escherichia coli
Davies, T. J.; Swann, J.; Sheppard, A. E.; Pickford, H.; Lipworth, S.; AbuOun, M.; Ellington, M.; Fowler, P. W.; Hopkins, S.; Hopkins, K.; Crook, D.; Peto, T. E.; Anjum, M. F.; Walker, A. S.; Stoesser, N.
Show abstract
2.Several bioinformatics genotyping algorithms are now commonly used to characterise antimicrobial resistance (AMR) gene profiles in whole genome sequencing (WGS) data, with a view to understanding AMR epidemiology and developing resistance prediction workflows using WGS in clinical settings. Accurately evaluating AMR in Enterobacterales, particularly Escherichia coli, is of major importance, because this is a common pathogen. However, robust comparisons of different genotyping approaches on relevant simulated and large real-life WGS datasets are lacking. Here, we used both simulated datasets and a large set of real E. coli WGS data (n=1818 isolates) to systematically investigate genotyping methods in greater detail. Simulated constructs and real sequences were processed using four different bioinformatic programs (ABRicate, ARIBA, KmerResistance, and SRST2, run with the ResFinder database) and their outputs compared. For simulations tests where 3,079 AMR gene variants were inserted into random sequence constructs, KmerResistance was correct for 3,076 (99.9%) simulations, ABRicate for 3,054 (99.2%), ARIBA for 2,783 (90.4%) and SRST2 for 2,108 (68.5%). For simulations tests where two closely related gene variants were inserted into random sequence constructs, KmerResistance the correct alleles in 35,338/46,318 (76.3%) ABRicate identified in 11,842/46,318 (25.6%) of simulations, ARIBA in 1679/46,318 (3.6%), and SRST2 in 2000/46,318 (4.3%). In real data, across all methods, 1392/1818 (76%) isolates had discrepant allele calls for at least one gene. Our evaluations revealed poor performance in scenarios that would be expected to be challenging (e.g. identification of AMR genes at <10x coverage, discriminating between closely related AMR gene sequences), but also identified systematic sequence classification (i.e. naming) errors even in straightforward circumstances, which contributed to 1081/3092 (35%) errors in our most simple simulations and at least 2530/4321 (59%) discrepancies in real data. Further, many of the remaining discrepancies were likely "artefactual" with reporting cut-off differences accounting for at least 1430/4321 (33%) discrepants. Comparing outputs generated by running multiple algorithms on the same dataset can help identify and resolve these artefacts, but ideally new and more robust genotyping algorithms are needed. 3. Impact statementWhole-genome sequencing is widely used for studying the epidemiology of antimicrobial resistance (AMR) genes in bacteria; however, there is some concern that outputs are highly dependent on the bioinformatics methods used. This work evaluates these concerns in detail by comparing four different, commonly used AMR gene typing methods using large simulated and real datasets. The results highlight performance issues for most methods in at least one of several simulated and real-life scenarios. However most discrepancies between methods were due to differential labelling of the same sequences related to the assumptions made regarding the underlying structure of the reference resistance gene database used (i.e. that resistance genes can be easily classified in well-defined groups). This study represents a major advance in quantifying and evaluating the nature of discrepancies between outputs of different AMR typing algorithms, with relevance for historic and future work using these algorithms. Some of the discrepancies can be resolved by choosing methods with fewer assumptions about the reference AMR gene database and manually resolving outputs generated using multiple programs. However, ideally new and better methods are needed.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Discordant bioinformatic predictions of antimicrobial resistance from whole-genome sequencing data of bacterial isolates: An inter-laboratory study 96%
- Diverse Genetic Determinants of Nitrofurantoin Resistance in UK Escherichia coli 96%
- K-mer based prediction of Clostridioides difficile relatedness and ribotypes 95%
Similar papers in this journal
- Accurate and Reproducible Whole-Genome Genotyping for Bacterial Genomic Surveillance with Nanopore Sequencing Data 96%
- Hash-based core genome multi-locus sequencing typing for Clostridium difficile 95%
- Population genomic molecular epidemiological study of macrolide resistant Streptococcus pyogenes in Iceland,1995-2016: Identification of a large clonal population with a pbp2x mutation conferring reduced in vitro beta-lactam susceptibility 94%
Similar papers in this journal
- Rapid nanopore metagenomic sequencing and predictive susceptibility testing of positive blood cultures from intensive care patients with sepsis 95%
- Development of an amplicon nanopore sequencing strategy for detection of mutations conferring intermediate resistance to vancomycin in Staphylococcus aureus strains 95%
- Real-time Plasmid Transmission Detection Pipeline 95%
Similar papers in this journal
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 95%
- An accurate and interpretable model for antimicrobial resistance in pathogenic Escherichia coli from livestock and companion animal species 94%
- Characterization of genetic diversity and population structure within Staphylococcus chromogenes by multilocus sequence Typing 93%
Similar papers in this journal
- A Panel of Diverse Pseudomonas aeruginosa Clinical Isolates for Research and Development 94%
- Tracking Antimicrobial Resistant Organisms Timely (TAROT): A Workflow Validation Study for Successive Core-genome SNP-based Nosocomial Transmission Analysis 94%
- Subpopulations in clinical samples of M. tuberculosis can give rise to rifampicin resistance and shed light on how resistance is acquired 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.