Back

Early detection of ampicillin susceptibility in Enterococcus faecium with MALDI-TOF MS and machine learning

Pichl, T.; Miranda, L.; Wantia, N.; Borgwardt, K.; Sattler, J.

2025-07-14 microbiology
10.1101/2025.07.14.664787 bioRxiv
Show abstract

BackgroundEnterococcus faecium can cause severe infections and is often resistant to the first-line antibiotic ampicillin. Consequently, clinicians usually prescribe broad-spectrum antibiotics, promoting the selection of multidrug-resistant bacteria. In this study, we investigate the application of machine learning techniques to detect ampicillin susceptibility directly from MALDI-TOF mass spectrometry. This technique could enable an earlier optimised treatment in infections with ampicillin-susceptible E. faecium. MethodsTwo datasets of clinical E. faecium MALDI-TOF spectra and their resistance phenotype were analysed: our own Technical University of Munich (TUM) dataset and the publicly available MS-UMG dataset. We tested logistic regression (LR) and LightGBM models on each dataset via nested cross-validation and explored transferability on the respective other dataset. ResultsLightGBM demonstrated slightly better performance than LR in identifying susceptible isolates in the TUM dataset (area under the precision-recall curve (AUPRC) 0.907 {+/-} 0.016 vs 0.902 {+/-} 0.030) as well as in the MS-UMG dataset (AUPRC 0.902 {+/-} 0.029 vs 0.899 {+/-} 0.054). External validation demonstrated good model transferability (AUPRC of 0.784 {+/-} 0.039 when trained on MS-UMG; 0.804 {+/-} 0.013 when trained on TUM). SHAP analysis consistently identified a top-ranked spectral feature corresponding to a peak at an m/z of 5091 in resistant isolate spectra. ConclusionThis study demonstrates that LR and LightGBM models can identify ampicillin-susceptible E. faecium isolates from MALDI-TOF spectra and generalise well to unseen datasets. While clinical implementation currently still requires confirmatory testing, the addition of larger datasets in the future will support the development of more robust machine learning models.

Published in Journal of Global Antimicrobial Resistance · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.