Back

Neural network-based predictions of antimicrobial resistance in Salmonella spp. using k-mers counting from whole-genome sequences

Caniu, C. J.

2021-08-11 genomics
10.1101/2021.08.10.455825 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWArtificial intelligence-based predictions have emerged as a friendly and reliable tool for the surveillance of the antimicrobial resistance (AMR) worldwide. In this regard, genome databases typically include whole-genome sequencing (WGS) data containing AMR meta-data that can be used to train machine learning (ML) models, in order to predict phenotype features from genome samples. In this study, using a Neural Network (NN) architecture and the SGD-ADAM algorithm, we build ML antibiotic resistance models that can predict Minimum Inhibitory Concentrations (MICs) and antimicrobial susceptibility profiles of Salmonella spp. Data analysis was based on 7,268 genomes publicly available in PATRIC database, containing about 75,000 AMR annotations. ML models were built using reference-free k-mer analysis of whole-genome sequences, MIC measurements and susceptibility categories, obtaining robust and accurate results for 9 antibiotics belonging to beta-lactam, fluoroquinolone, phenicol, aminoglycoside, tetracycline and sulphonamide classes. Al-though the accuracy of predicting the actual MIC reaches modest levels, the within {+/-} 1 2-fold dilution accuracy per antibiotic reaches significant levels with values that varies from 85% to 95%, with narrow 95% CIs of about 5% and individual accuracies per MIC {gtrsim} 80%. For differentiation between "susceptible" and "resistant" values, by measuring the accuracy and error of models susceptibility predictions to different antibiotics, the accuracy is the same as before and ranges from 85% to 95%, with 95% CIs of about 5%, the recall extends from 75% to 85%, the precision from 60% to 90%, whereas the very major error is [lsim] 20%. In summary, these results show that NN-based models are able to learn and predict the AMR phenotype from bacterial genomes based on a gene-free k-mer analysis.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.