Back

Chemical Space Localization for Unknown Metabolite Annotation via Semantic Similarity of Mass Spectral Language

Ji, H.; Du, R.; Dai, Q.; Su, M.; Lyu, Y.; Yan, J.

2024-10-15 bioinformatics
10.1101/2024.05.30.596727 bioRxiv
Show abstract

Untargeted analysis using liquid chromatography{square}mass spectrometry (LC-MS) allows quantification of known and unknown compounds within biological systems. However, in practical analysis of complex biological system, the majority of compounds often remain unidentified. Here, we developed a novel deep learning-based compound annotation approach via semantic similarity analysis of mass spectral language. This approach enables the prediction of structurally related compounds for unknowns. By considering the chemical space, these structurally related compounds provide valuable information about the potential location of the unknown compounds and assist in ranking candidates obtained from molecular structure databases. Validated with two independent benchmark datasets obtained by chemical standards, our method has consistently demonstrated superior performance compared to existing compound annotation methods. A case study of the tomato ripening process indicates that DeepMASS has significant potential for metabolic biomarker identification in real biological systems. Overall, the presented method shows considerable promise in annotating metabolites, particularly in revealing the "dark matter" in untargeted analysis.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.