Reaction-Conditioned Enzyme Discovery with Multimodal Deep Learning
Tan, P.
Show abstract
The precise mapping between chemical transformations and enzymatic catalysts underpins the complexity of metabolic networks. Conventional discovery methods, tethered to sequence homology or structural alignment, are inherently blind to new reactions. Here we present VenusRXN, a multimodal deep learning framework that shatters this limitation by enabling reaction-conditioned enzyme discovery. By seamlessly unifying a pre-trained reaction encoder with a protein language model, VenusRXN achieves a fine-grained, high-dimensional alignment of chemical and biological representations. On benchmarks to discover enzymes which catalyze reactions not seen in the training dataset, it surpasses state-of-the-art baselines with a top-20 retrieval hit rate of 76.5%. Most critically, we demonstrate VenusRXNs capabilities in a zero-shot discovery. As verified by the wet-lab experiments, it successfully identified enzymes to catalyze the chemical reactions never reported, including the one to catalyze the synthesis route for a type 2 diabetes drug intermediate using a non-natural substrate. With surprising precision, the model pinpointed active candidates within the top 10 sequences directly from a global search space of over 300 million proteins, which can hardly be achieved by structure-based enzyme discovery algorithm. Thus, VenusRXN unlocks the capacity to interrogate the vast, unanno-tated "dark matter" of the protein universe with affordable computational cost for everyone. This work signals a definitive paradigm shift, establishing the chemical reaction itself, rather than homology, as the primary functional descriptor for the de novo discovery of biocatalysts.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CLOOME: contrastive learning unlocks bioimaging databases for queries with chemical structures 95%
- scDisInFact: disentangled learning for integration and prediction of multi-batch multi-condition single-cell RNA-sequencing data 95%
- Protein language model powers accurate and fast sequence search for remote homology 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.