Back

IDBac: an open-access web platform and compendium for the identification of bacteria by MALDI-TOF mass spectrometry.

Krull, N. K.; Strobel, M.; Saulog, J.; Zaroubi, L.; Paulo, B. S.; Timba, M.; Braun, D. R.; Mingolelli, G.; Raherisoanjato, J.; Shepherd, R. A.; Scott, A. F.; De Silva, C.; Fergusson, C.; Daniel, Z.; Pokharel, S. K.; Romanowski, S.; Hernandez, A.; Monge-Loria, M.; Dylla, C. E.; Natu, M. M.; Petukhova, V. Z.; Garg, N.; Jensen, P. R.; Blachowicz, A.; Cassilly, C. D.; Guan, L.; Stevens, C. D.; Winter, J. M.; McKinnie, S. M. K.; Adaikpoh, B. I.; Carlson, S.; McCauley, E. P.; Metcalf, W. W.; Bugni, T. S.; Mullowney, M. W.; Pamer, E. G.; Henke, M. T.; Barton, H.; Carter, D. O.; Eustaquio, A. S.; Lini

2025-10-15 microbiology
10.1101/2025.10.15.682631 bioRxiv
Show abstract

The identification and analysis of bacteria is central to the microbiological sciences. While gene sequencing methods have been the standard to achieve this, use of MALDI-TOF mass spectrometry (MS), particularly in clinical microbiology, provides high-throughput identification to the subspecies level. However, biotyping has yet to be adopted outside of clinical settings due to the lack of a centralized public database of MS protein signatures that would facilitate isolate identification via spectral comparison. Further, current platforms lack meaningful ways to compare multiple properties from large numbers of bacterial isolates. Herein we present the IDBac web platform, a crowd-sourced central knowledgebase of protein MS signatures of >1400 strains spanning 6 bacterial phyla. Accompanying the knowledgebase is analysis infrastructure to identify unknown isolates, probe relationships within culture collections using metadata integration, and visualize specialized metabolite differences within groups of closely related bacteria. To highlight this utility and encourage wide community contribution, examples of each are presented.

Published in Nature Communications (predicted rank #2) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.