Back

Reaction-Conditioned Enzyme Discovery with Multimodal Deep Learning

Tan, P.

2026-03-10 synthetic biology
10.64898/2026.03.09.710689 bioRxiv
Show abstract

The precise mapping between chemical transformations and enzymatic catalysts underpins the complexity of metabolic networks. Conventional discovery methods, tethered to sequence homology or structural alignment, are inherently blind to new reactions. Here we present VenusRXN, a multimodal deep learning framework that shatters this limitation by enabling reaction-conditioned enzyme discovery. By seamlessly unifying a pre-trained reaction encoder with a protein language model, VenusRXN achieves a fine-grained, high-dimensional alignment of chemical and biological representations. On benchmarks to discover enzymes which catalyze reactions not seen in the training dataset, it surpasses state-of-the-art baselines with a top-20 retrieval hit rate of 76.5%. Most critically, we demonstrate VenusRXNs capabilities in a zero-shot discovery. As verified by the wet-lab experiments, it successfully identified enzymes to catalyze the chemical reactions never reported, including the one to catalyze the synthesis route for a type 2 diabetes drug intermediate using a non-natural substrate. With surprising precision, the model pinpointed active candidates within the top 10 sequences directly from a global search space of over 300 million proteins, which can hardly be achieved by structure-based enzyme discovery algorithm. Thus, VenusRXN unlocks the capacity to interrogate the vast, unanno-tated "dark matter" of the protein universe with affordable computational cost for everyone. This work signals a definitive paradigm shift, establishing the chemical reaction itself, rather than homology, as the primary functional descriptor for the de novo discovery of biocatalysts.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.