Development and Validation of a MALDI-TOF-Based Model to Predict Extended-Spectrum Beta-Lactamase and/or Carbapenemase-Producing in Klebsiella pneumoniae Clinical Isolates
Guerrero-Lopez, A.; Candela, A.; Sevilla-Salcedo, C.; Hernandez-Garcia, M.; Martinez-Olmos, P.; Canton, R.; Munoz, P.; del Campo, R.; Gomez-Verdejo, V.; Rodriguez-Sanchez, B.
Show abstract
Matrix-Assisted Laser Desorption Ionization Time-Of-Flight (MALDI-TOF) Mass Spectrometry (MS) is a reference method for microbial identification and it can be used to predict Antibiotic Resistance (AR) when combined with artificial intelligence methods. However, current solutions need time-costly preprocessing steps, are difficult to reproduce due to hyperparameter tuning, are hardly interpretable, and do not pay attention to epidemiological differences inherent to data coming from different centres, which can be critical. We propose using a multi-view heterogeneous Bayesian model (KSSHIBA) for the prediction of AR using MALDI-TOF MS data together with their epidemiological differences. KSSHIBA is the first model that removes the ad-hoc preprocessing steps that work with raw MALDI-TOF data. In addition, due to its Bayesian probabilistic nature, it does not require hyperparameter tuning, provides interpretable results, and allows exploiting local epidemiological differences between data sources. To test the proposal, we used data from 402 Klebsiella pneumoniae isolates coming from two different domains and 20 different hospitals located in Spain and Portugal. KSSHIBA outperforms current state-of-the-art approaches in antibiotic susceptibility prediction, obtaining a 0.78 AUC score in Wild Type classification and a 0.90 AUC score in Extended-Spectrum Beta-Lactamases (ESBL)+Carbapenemases (CP)-producers. The proposal consistently removes the need for ad-hoc preprocessing by working with raw MALDI-TOF data, which, in turn, reduces the time needed to obtain the results of the resistance mechanism in microbiological laboratories. The proposed model implementation as well as both data domains are publicly available.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Ranking microbial metabolomic and genomic links in the NPLinker framework using complementary scoring functions 93%
- SPARTA: Interpretable functional classification of microbiomes and detection of hidden cumulative effects. 93%
- iPRESTO: automated discovery of biosynthetic sub-clusters linked to specific natural product substructures 92%
Similar papers in this journal
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 94%
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 93%
- Comparative Performance Of Two Automated Machine Learning Platforms For COVID-19 Detection By Maldi-Tof-Ms 92%
Similar papers in this journal
- Feature selection with vector-symbolic architectures: a case study on microbial profiles of shotgun metagenomic samples of colorectal cancer 95%
- Comparative analysis of machine learning algorithms on the microbial strain-specific AMP prediction 95%
- Applied Machine Learning for human bacteriaMALDI-TOF Mass Spectrometry: a systematicreview 94%
Similar papers in this journal
- PathoGFAIR: a collection of FAIR and adaptable (meta)genomics workflows for (foodborne) pathogens detection and tracking 93%
- CoCoPyE: feature engineering for learning and prediction of genome quality indices 92%
- IDseq - An Open Source Cloud-based Pipeline and Analysis Service for Metagenomic Pathogen Detection and Monitoring 91%
Similar papers in this journal
- Using Deep Learning for Gene Detection and Classification in Raw Nanopore Signals 93%
- DAnIEL: A User-Friendly Web Server for Fungal ITS Amplicon Sequencing Data 92%
- BacAnt: A Combination Annotation Server for Bacterial DNA Sequences to Identify Antibiotic Resistance Genes, Integrons, and Transposable Elements. 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.