Entropy-based decoy generation methods for accurate FDR estimation in large-scale metabolomics annotations.
An, S.; Lu, M.; Wang, R.; Wang, J.; Xie, C.; Tong, J.; Jiang, H.; Yu, C.
Show abstract
Large-scale metabolomics research faces challenges in accurate metabolite annotation and false discovery rate (FDR) estimation. Recent progress in addressing these challenges has leveraged experience from proteomics and inspiration from other sciences. Although the target-decoy strategy has been applied to metabolomics, generating reliable decoy libraries is difficult due to the complexity of metabolites. Additionally, continuous bioinformatic efforts are necessary to increase the utilization of growing spectra resources while reducing false identifications. Here we introduce the concept of ion entropy and present two entropy-based decoy generation methods. The assessment of public spectral databases using ion entropy validated it as a good metric for ion information content in massive metabolomics data. The decoy generation method developed based on this concept outperformed current representative decoy strategies in metabolomics and achieved the best FDR estimation performance. We analyzed 47 public metabolomics datasets using the constructed workflow to provide instructive suggestions. Finally, we present MetaPhoenix, a tool equipped with a well-constructed FDR estimation workflow that facilitates the development of accurate FDR-controlled analysis in the metabolomics field.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- mzrtsim: Raw Data Simulation for Reproducible Gas/Liquid Chromatography Mass Spectrometry Based Non-targeted Metabolomics Data Analysis 97%
- spectrum_utils: A Python package for mass spectrometry data processing and visualization 96%
- A Full Window Data Independent Acquisition Method for DeeperTop-down Proteomics 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MMS2plot: an R package for visualizing multiple MS/MS spectra for groups of modified and non-modified peptides 97%
- DrawAlignR: An interactive tool for across run chromatogram alignment visualization 95%
- Leveraging immonium ions for identifying and targeting acyl-lysine modifications in proteomic datasets 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.