Back

Ontology-guided harmonization enables unified discovery of public metabolomics studies within and across repositories

Banerjee, S.; Jalan, P.; Chinhara, R.; Kalle, C.; Wangikar, P.; Jadhav, K.

2026-07-27 bioinformatics
10.64898/2026.07.23.740366 bioRxiv
Show abstract

Public metabolomics repositories contain thousands of studies, but differences in metadata structure, vocabulary, and repository-specific terms still limit reliable search, comparison, and reuse within and across databases. Here we present HARMONY, an ontology-based framework and web platform that harmonizes study-level metadata and metabolite information across Metabolomics Workbench and MetaboLights studies. HARMONY resolves eight biological and analytical metadata nodes, including species, sample source, disease, analytical technique, separation method, ion polarity, ionization source, and mass analyzer type, while preserving the original deposited terms as evidence. A ninth node, metabolite identity, maps metabolite entities to RefMet across both repositories. HARMONY uses a two-step workflow: Multi-source extraction retrieves records missed by single-field lookups, and ontology mapping then converts repository-specific labels into shared query terms, substantially closing the cross-repository retrieval gap relative to raw matching. Across the full corpus, HARMONY increased cross-repository retrievability from 75.5% to 89.6%, yielding thousands of study-node retrievals and reconnecting studies that raw-text search would have left unreachable within their own repositories. Approximately 91% of Metabolomics Workbench and 85% of MetaboLights studies had at least six of the eight nodes harmonized. The resulting platform, available at https://omicsinharmony.in, supports ontology-aware search, metadata filtering, within- and cross-repository study comparison, and metabolite-level querying, with retrieval backed by machine learning encoders that map study metadata into shared representations of biological and analytical context. HARMONY provides the metabolomics community with a shared, traceable search interface for study discovery and comparison within and across public repositories.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.