IDSL.UFA assigns high confidence molecular formula annotations for untargeted LC/HRMS datasets in metabolomics and exposomics
Fakouri-Baygi, S.; Banerjee, S.; Chakraborty, P.; Kumar, Y.; Barupal, D. K.
Show abstract
Untargeted LC/HRMS assays in metabolomics and exposomics aim to characterize the small molecule chemical space in a biospecimen. To gain maximum biological insights from these datasets, LC/HRMS peaks should be annotated with chemical and functional information including molecular formula, structure, chemical class and metabolic pathways. Among these, molecular formulas may be assigned to LC/HRMS peaks through matching theoretical and observed isotopic profiles (MS1) of the underlying ionized compound. For this, we have developed the Integrated Data Science Laboratory for Metabolomics and Exposomics - United Formula Annotation (IDSL.UFA) R package. In the untargeted metabolomics validation tests, IDSL.UFA assigned 54.31%-85.51% molecular formula for true positive annotations as the top hit, and 90.58%-100% within the top five hits. Molecular formula annotations were also supported by MS/MS data. We have implemented new strategies to 1) generate formula sources and their theoretical isotopic profiles 2) optimize the formula hits ranking for the individual and the aligned peak lists and 3) scale IDSL.UFA-based workflows for studies with larger sample sizes. Annotating the raw data for a publicly available pregnancy metabolome study using IDSL.UFA highlighted hundreds of new pregnancy related compounds, and also suggested presence of chlorinated perfluorotriether alcohols (Cl-PFTrEAs) in human specimens. IDSL.UFA is useful for human metabolomics and exposomics studies where we need to minimize the loss of biological insights in untargeted LC/HRMS datasets. The IDSL.UFA package is available in the R CRAN repository https://cran.r-project.org/package=IDSL.UFA. Detailed documentation and tutorials are also provided at www.ufa.idsl.me.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Introducing identification probability for automated and transferable assessment of metabolite identification confidence in metabolomics and related studies 98%
- Met-ID: An Open-Source Software for Comprehensive Annotation of Multiple On-Tissue Chemical Modifications in MALDI-MSI 97%
- MS-CleanR: A feature-filtering approach to improve annotation rate in untargeted LC-MS based metabolomics 96%
Similar papers in this journal
- ModiFinder: Tandem Mass Spectral Alignment Enables Structural Modification Site Localization 96%
- Using variable data independent acquisition for capillary electrophoresis-based untargeted metabolomics 96%
- Tracking the Metabolic Fate of Exogenous Arachidonic Acid in Ferroptosis Using Dual-Isotope Labeling Lipidomics 96%
Similar papers in this journal
- Public LC-Orbitrap-MS/MS Spectral Library for Metabolite Identification 96%
- Development and Application of Multidimensional Lipid Libraries to Investigate Lipidomic Dysregulation Related to Smoke Inhalation Injury Severity 96%
- Hybrid Quadrupole Mass Filter Radial Ejection Linear Ion Trap and Intelligent Data Acquisition Enable Highly Multiplex Targeted Proteomics 95%
Similar papers in this journal
- MS2Lipid: a lipid subclass prediction program using machine learning and curated tandem mass spectral data 96%
- High-throughput UHPLC-MS to screen metabolites in feces for gut metabolic health 96%
- Scan-Centric, Frequency-Based Method for Characterizing Peaks from Direct Injection Fourier transform Mass Spectrometry Experiments 95%
Similar papers in this journal
- A comprehensive LFQ benchmark dataset on modern day acquisition strategies in proteomics 94%
- Implementing the re-use of public DIA proteomics datasets: from the PRIDE database to Expression Atlas 94%
- An interactive mass spectrometry atlas of histone posttranslational modifications in T-cell acute leukemia 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.