Back

Predicting Enzyme pH Optima from Structure Using Equivariant Graph Neural Networks

SinhaRoy, R.; Clauss, C.; Ivanikov, I.; Kuenze, G.

2026-01-21 bioinformatics
10.64898/2026.01.18.700076 bioRxiv
Show abstract

Enzyme activity and stability are strongly modulated by pH, making the catalytic pH optimum (pH opt) a key parameter in enzyme development and biotechnological applications. Experimental determination of pH opt is, however, labor-intensive and time-consuming, motivating the development of accurate computational prediction methods. Here, we introduce pHoptNN, an E (n)-equivariant graph neural network designed to predict enzyme pH opt directly from three-dimensional protein structures. pHoptNN was trained on a curated dataset comprising nearly 12,000 enzymes with experimentally determined pH opt values and high-confidence structural models obtained from the Protein Data Bank and AlphaFold3. The model represents enzymes as atomic-level molecular graphs, integrating structural, chemical, and electrostatic features. Model development was guided by extensive hyperparameter optimization using genetic and Bayesian search strategies. On a held-out test set, pHoptNN achieved a root-mean-square error (RMSE) of 0.588 pH units, substantially outperforming the sequence-based method EpHod (RMSE = 0.879). Moreover, pHoptNN maintains robust predictive performance across different enzyme classes and pH ranges. These results demonstrate the utility of structure-based equivariant deep learning for enzyme pH opt prediction and highlight the potential of pHoptNN to accelerate enzyme discovery and engineering workflows.

Published in Journal of Chemical Informationand Modeling · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.