Quantifying per-match Reliability in Library Matching for Untargeted Metabolomics Workflows
Charria-Giron, E.; van IJcken, J.; Della Vedova, L.; Torres-Ortega, L. R.; van der Hooft, J. J. J.
Show abstract
Tandem mass spectrometry has become central to untargeted metabolomics. The translation of unknown spectra into biological insight depends on assigning chemical identities to detected metabolites. Structural characterization typically begins with mass spectral library matching, in which experimental spectra are compared against reference libraries and candidate annotations are ranked by their spectral similarity to the query. As spectral libraries and experimental datasets grow, however, more candidates achieve comparable similarity scores for a single query, and similarity scores give no indication of how reproducible a candidate match is or how sensitive it is to the underlying fragment evidence. Existing false-discovery-rate approaches can indicate annotation error at the dataset level but do not provide a per-match estimate of reliability. Here, we introduce a SpecReBoot-inspired query-focused bootstrapping approach that resamples the fragment evidence of each query spectrum. This approach relies on recomputing query similarity to candidate library spectra across bootstrap replicates, which provides a statistical distribution of scores rather than a single value. From this distribution we define the match support, a per-match reliability estimate quantifying the reproducibility of a match under spectral perturbation, together with measures of ranking stability that describe how often a candidate remains among the top-ranked matches across replicates. Applied to a forensic drug-of-abuse case, match support distinguished previously identified annotations from high-scoring false positives: a distinction cosine similarity failed to make. Furthermore, match support values remained stable as the reference library was expanded, whereas ranking stability metrics shifted significantly. In a cross-instrument endogenous metabolite library search, match support further revealed metric-specific annotation behavior, identifying metabolites consistently supported across different similarity metrics, while flagging annotations whose reliability depended strongly on the chosen scoring metric. Benchmarking against a natural-product reference library demonstrated that ranking based on match support values promoted true matches by four ranks on average compared with cosine-based ranking, without promoting analogs. Under controlled spectral perturbation experiments, match support flagged incorrect annotations with an AUROC of 0.75, whereas the cosine similarity score alone of the same match reached only 0.56. Query-focused bootstrapping thus provides a practical, per-match measure of annotation reliability, bringing the field a step toward reliable annotations at scale. We anticipate that incorporation of our annotation reliability scoring into computational metabolomics workflows will further promote the growth of spectral libraries and enhance their applicability across scientific disciplines.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AutoTuner: High fidelity, robust, and rapid parameter selection for metabolomics data processing 97%
- rtmsEcho: An Open-Source R Package for Automated Analysis of Acoustic Ejection Mass Spectrometry Data 96%
- Introducing identification probability for automated and transferable assessment of metabolite identification confidence in metabolomics and related studies 96%
Similar papers in this journal
- ModiFinder: Tandem Mass Spectral Alignment Enables Structural Modification Site Localization 97%
- High-Throughput Measurement and Machine Learning-Based Prediction of Collision Cross Sections for Drugs and Drug Metabolites 96%
- Comparison of Cosine, Modified Cosine, and Neutral Loss Based Spectrum Alignment For Discovery of Structurally Related Molecules 95%
Similar papers in this journal
- Inserting Pre-Analytical Chromatographic Priming Runs Significantly Improves Targeted Pathway Proteomics With Sample Multiplexing 95%
- Scribe: next-generation library searching for DDA experiments 95%
- Increasing the Throughput and Reproducibility of Activity-Based Proteome Profiling Studies with Hyperplexing and Intelligent Data Acquisition 95%
Similar papers in this journal
- A New Platform for Label-Free, Proximal Cellular Pharmacodynamic Assays: Identification of Glutaminase Inhibitors Using Infrared Matrix-Assisted Laser Desorption Electrospray Ionization Mass Spectrometry 93%
- Discovery of a celecoxib binding site on PTGES with a cleavable chelation-assisted biotin probe 93%
- Dual-Probe Activity-Based Protein Profiling Reveals Site-Specific Differences in Protein Binding of EGFR-Directed Drugs 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.