Back

METEOR: joint genome-scale reconstruction and enzyme prediction

Niu, K.; Kulmanov, M.; Hoehndorf, R.

2026-02-02 bioinformatics
10.64898/2026.02.01.701942 bioRxiv
Show abstract

MotivationCurrent machine learning methods for enzyme function prediction primarily treat proteins as independent entities, ignoring the metabolic context in which they operate. This reductionist approach often generates biologically implausible annotations that fail to satisfy stoichiometric or thermodynamic constraints. While genome-scale metabolic models (GEMs) enforce systemic coherence, traditional reconstruction workflows decouple functional annotation from network assembly, resulting in information loss during the discretization of enzymatic evidence. ResultsWe developed METEOR, a framework that integrates continuous confidence scores from deep learning models directly into a Mixed-Integer Linear Programming (MILP) formulation. By conditioning predictions on global metabolic requirements, METEOR resolves functional ambiguities where sequence signals are weak. Evaluation on the Price-149 dataset shows that our method yields physiologically viable networks capable of sustaining growth and improve protein-level predictions over baseline methods. Validation on 6,894 bacterial genomes reveals that system-aware refinement increases the recall of experimentally observed phenotypes substantially. Furthermore, we show that constraint-driven selection can still improve annotation performance in highly incomplete genomes. Our results suggest that functional annotation should be treated as a unified inference problem where global system constraints supervise local predictions. AvailabilityMETEOR is freely available at https://github.com/bio-ontology-research-group/Meteor Contactrobert.hoehndorf@kaust.edu.sa Supplementary informationSupplementary data are available at Bioinformatics online.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.