PocketGNN: A Cross-Modal Framework Unifying Local 3D Pocket Geometry and Global Sequence Semantics for Enzyme Kinetic Prediction
Li, Z.; Lu, D.
Show abstract
The enzyme turnover number (kcat) is a pivotal kinetic parameter for understanding bio-catalytic efficiency, yet its accurate prediction remains a grand challenge due to the complex interplay between local physicochemical constraints and global evolutionary context. Existing methods typically bifurcate into sequence-based approaches, which capture evolutionary semantics but miss fine-grained spatial details, or structure-based models, which often suffer from noise in whole-protein representations or lack global context. To bridge this gap, we propose PocketGNN, a cross-modal deep learning framework that synergizes the precision of local 3D geometry with the breadth of global 1D sequence semantics. PocketGNN introduces a high-fidelity graph representation of the active pocket, enriched with a novel 24-dimensional geometric edge encoding (RBF distances, bond angles, dihedral angles) to capture the stereochemical determinants of catalysis. Crucially, this local structural view is fused with global evolutionary information extracted from pre-trained protein language models (ESM-2), creating a unified representation that spans spatial scales. Evaluated on a rigorous dataset derived from IntEnzyDB, PocketGNN achieves a Pearson correlation coefficient (r) of 0.98 and a coefficient of determination (R2) of 0.918 for log10(kcat) under standard random splitting. Furthermore, under a strict 40% sequence identity split designed to test zero-shot generalization to unseen families, the model maintains a robust correlation (r = 0.67, R2 = 0.44), significantly outperforming recent state-of-the-art methods including CatPred (r = 0.52) and CataPro (r = 0.50). Interpretability analysis confirms that the model successfully attends to key catalytic residues, validating its ability to learn chemically meaningful structure-function relationships rather than mere sequence memorization.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 95%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.