Back

MSMICA: computational metabolite identification in untargeted metabolomics by integrating MS, retention time, and biological evidence

Zhan, J.; Weinberg, J.; Crandall, W. J.; Qin, Z.; Jarrell, Z. R.; Preston, J. D.; Nellis, M.; Teeny, S.; Liang, D.; Martin, G. S.; Price, N. L.; de Cabo, R.; Master, V.; Cohn, B. A.; Go, Y.-M.; Jones, D. P.

2026-08-20 bioinformatics
10.64898/2026.08.15.744986 bioRxiv
Show abstract

Mass Spectrometry Metabolomics Identification Connection Algorithm (MSMICA) is an algorithm for automated metabolite identification in untargeted liquid chromatography-high-resolution mass spectrometry (LC-HRMS) analyses. Limitations in metabolite identification can occur due to the availability and cost of standards and prevent recognition of metabolic factors impacting human health and disease. MSMICA performs mass-to-charge-ratio matching with chemical structures and clusters of LC-HRMS features for adduct and isotope forms. A local optimization is then used to integrate retention time prediction, metabolite precursor-product and transporter correlations, and biospecimen-specific abundance information for metabolite identification. Applying MSMICA to various internal and external mammalian datasets, validation results showed a 96.2 +- 5.1% correct rate of metabolite identification. When multiple LC-HRMS datasets were used, MSMICA enabled greater metabolite identifications, expanded metabolic pathway coverage, and data harmonization. Thus, MSMICA applies multiple pieces of evidence to substantially improve metabolite identification coverage and accuracy for known metabolites.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.