Data processing of product ion spectra: Methods to control false discovery rate in compound search results for non-targeted metabolomics
Matsuda, F.
Show abstract
In non-targeted metabolomics utilizing high-resolution mass spectrometry, several database search methods have been used to comprehensively annotate the acquired product ion spectra. Recent advancements in various in silico prediction techniques have facilitated compound searches by scoring the degree of coincidence between a query product ion spectrum and a compound in a compound database. Certain search results may be false positives, thus necessitating a method for controlling the false discovery rate (FDR). This study proposed two simple methods for controlling the FDR in compound search results. In the pseudo-target decoy method, the FDR can be estimated without creating a separate decoy database by treating such as the positive ion mode spectra as targets and converting the negative ion mode spectra as decoys. Further, the second-rank method uses the score distribution of the second-ranked hits from the compound search as an approximation of the false-positive distribution of the top-ranked hits. The performance of these methods was evaluated by annotating the product ion spectra from MassBank using the SIRIUS 5 CSI:Finger ID scoring method. The results indicated that the second-rank method was closer to the true FDR of 0.05. When applied to the four human metabolomics datasets, the second-rank method provided more conservative FDR estimations than the pseudo-target-decoy method. These methods enabled the identification of metabolites not present in human metabolome databases. Overall, this study demonstrates the utility of these simple methods for FDR control in non-targeted metabolomics, facilitating more reliable compound identification and the potential discovery of novel metabolites.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Met-ID: An Open-Source Software for Comprehensive Annotation of Multiple On-Tissue Chemical Modifications in MALDI-MSI 96%
- Using data-dependent and independent hybrid acquisitions for fast liquid chromatography-based untargeted lipidomics 96%
- TopLib: Building and searching top-down mass spectral libraries for proteoform identification 96%
Similar papers in this journal
- MAFFIN: Metabolomics Sample Normalization Using Maximal Density Fold Change with High-Quality Metabolic Features and Corrected Signal Intensities 96%
- Identification of metabolites from tandem mass spectra with a machine learning approach utilizing structural features 95%
- 3DMolMS: Prediction of Tandem Mass Spectra from Three Dimensional Molecular Conformations 95%
Similar papers in this journal
Similar papers in this journal
- Deep Learning-based Pseudo-Mass Spectrometry Imaging Analysis for Precision Medicine 95%
- SingleFrag: A deep learning tool for MS/MS fragment and spectral prediction and metabolite annotation 94%
- ChemEmbed: A deep learning framework for metabolite identification using enhanced MS/MS data and multidimensional molecular embeddings 93%
Similar papers in this journal
- MS2Lipid: a lipid subclass prediction program using machine learning and curated tandem mass spectral data 97%
- Matrix selection for the visualization of small molecules and lipids in brain tumors using untargeted MALDI-TOF mass spectrometry imaging 95%
- Robust Moiety Model Selection Using Mass Spectrometry Measured Isotopologues 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.