Back

Entropy-based decoy generation methods for accurate FDR estimation in large-scale metabolomics annotations.

An, S.; Lu, M.; Wang, R.; Wang, J.; Xie, C.; Tong, J.; Jiang, H.; Yu, C.

2023-07-02 bioinformatics
10.1101/2023.07.02.547371 bioRxiv
Show abstract

Large-scale metabolomics research faces challenges in accurate metabolite annotation and false discovery rate (FDR) estimation. Recent progress in addressing these challenges has leveraged experience from proteomics and inspiration from other sciences. Although the target-decoy strategy has been applied to metabolomics, generating reliable decoy libraries is difficult due to the complexity of metabolites. Additionally, continuous bioinformatic efforts are necessary to increase the utilization of growing spectra resources while reducing false identifications. Here we introduce the concept of ion entropy and present two entropy-based decoy generation methods. The assessment of public spectral databases using ion entropy validated it as a good metric for ion information content in massive metabolomics data. The decoy generation method developed based on this concept outperformed current representative decoy strategies in metabolomics and achieved the best FDR estimation performance. We analyzed 47 public metabolomics datasets using the constructed workflow to provide instructive suggestions. Finally, we present MetaPhoenix, a tool equipped with a well-constructed FDR estimation workflow that facilitates the development of accurate FDR-controlled analysis in the metabolomics field.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.