Back

Towards automatic derivation of geometry-based descriptors as surrogates for complex computational approaches in enzyme-substrate prediction

Sequeiros-Borja, C. E.; Skoda, P.; Brezovsky, J.

2025-12-01 bioinformatics
10.1101/2025.11.26.690723 bioRxiv
Show abstract

Accurate prediction of enzyme-substrate interactions remains a fundamental challenge in biocatalysis and drug discovery. While machine learning approaches have shown promise, they require extensive training data and often lack mechanistic interpretability. Here, we present a novel methodology that automatically derives geometry-based descriptors from enzyme-substrate complex structures to predict substrate specificity. Our approach simplifies complex catalytic mechanisms into interpretable geometric filters comprising critical inter-atomic distances and accessibility of atomic pairs parameters. We validated this methodology using two mechanistically distinct enzyme families with minimal training data: haloalkane dehalogenases (9 enzymes and 53 substrates) and aldehyde reductases (9 enzymes and 36 substrates). The filters demonstrated robust performance across chemically diverse substrates. On testing datasets, the derived filters achieved average accuracy of 77% and sensitivity of 94% for haloalkane dehalogenases and average 57% recall of true substrates for aldehyde reductases, exceeding state-of-the-art machine learning methods for substrate predictions on these datasets. Crucially, the geometric descriptors directly correspond to catalytic requirements, providing mechanistic insights into substrate recognition. This interpretable, mechanism-based approach requires minimal training data and can be readily applied to newly characterized enzymes, offering a powerful tool for enzyme engineering and substrate screening applications.

Published in ChemPhysChem · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.