Back

MAERM: Predicting Enzyme-Reaction Matching Relationships with a Mixed-Attention Model

Liu, T.; Zhai, S.; Lin, S.; Zhan, X.; Deng, J.; Liu, H.; Siu, S. W. I.

2026-07-10 bioinformatics
10.64898/2026.07.06.736902 bioRxiv
Show abstract

Harnessing enzyme specificity requires a thorough understanding of enzyme promiscuity, which determines enzymes catalytic scope; however, measuring this scope still relies heavily on labor-intensive analytical approaches. While data-driven approaches have emerged to predict the catalytic scope of enzymes, these methods continue to face challenges such as restricted datasets and insufficient integration of enzyme structural information and reaction transformations. Here, we introduce MAERM, an innovative mixed-attention model designed to predict enzyme-reaction matching relationships. Built on our MAERM-DB, a dataset with broad coverage of validated and chemoenzymatic catalysis data, MAERM utilizes a local-global attention module to integrate multimodal enzyme information with fine-grained reaction representations, thereby predicting enzyme-reaction matching probabilities. Results show that MAERM consistently outperforms all baselines, with an average F1-score of 0.984. Notably, on challenging test samples with less than 40% sequence identity to the training set, MAERM outperforms the second-ranked model by 5.9% in F1-score. In addition, MAERM achieves the highest top-10 success rate of 51.7% on Enzyme-405 and the highest balanced accuracy of 0.697 on BioCat-547, further supporting its generalizability in enzyme screening and chemoenzymatic catalysis. Finally, MAERM can serve as an efficient scoring module. When integrated with ProteinMPNN, MAERM has successfully guided novel enzyme design for two carbonyl reduction reactions, resulting in enhanced catalytic potential for the native substrate and demonstrating broad compatibility. Overall, MAERM has the potential to reduce the experimental cost of measuring enzymes catalytic scope, facilitate enzyme design, and ultimately accelerate the design-build-test-learn cycle in enzyme engineering.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

1
Nature Communications
5641 papers in training set
Top 12%
14.8%
2
Journal of Chemical Information and Modeling
238 papers in training set
Top 0.5%
11.7%
3
Communications Chemistry
48 papers in training set
Top 0.1%
7.8%
4
ACS Catalysis
18 papers in training set
Top 0.1%
5.4%
5
Journal of Chemical Theory and Computation
140 papers in training set
Top 0.4%
4.3%
6
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.0%
7
Nature Machine Intelligence
70 papers in training set
Top 0.7%
4.0%
50% of probability mass above
8
Chemical Science
73 papers in training set
Top 0.5%
3.2%
9
Computational and Structural Biotechnology Journal
242 papers in training set
Top 2%
2.4%
10
Protein Science
246 papers in training set
Top 2%
2.4%
11
Advanced Science
286 papers in training set
Top 3%
2.4%
12
ACS Central Science
71 papers in training set
Top 0.6%
2.1%
13
PLOS Computational Biology
1863 papers in training set
Top 13%
2.1%
14
Nature Chemical Biology
119 papers in training set
Top 1%
1.9%
15
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 28%
1.7%
16
Bioinformatics
1204 papers in training set
Top 7%
1.7%
17
Journal of Cheminformatics
29 papers in training set
Top 0.5%
1.5%
18
Nucleic Acids Research
1281 papers in training set
Top 11%
1.1%
19
JACS Au
43 papers in training set
Top 0.7%
1.0%
20
Communications Biology
993 papers in training set
Top 25%
1.0%
21
Cell Systems
201 papers in training set
Top 4%
1.0%
22
ACS Synthetic Biology
287 papers in training set
Top 2%
1.0%
23
Journal of Molecular Biology
232 papers in training set
Top 3%
1.0%
24
ChemBioChem
55 papers in training set
Top 1.0%
1.0%
25
Scientific Reports
3612 papers in training set
Top 71%
1.0%
26
ACS Omega
105 papers in training set
Top 3%
1.0%
27
Synthetic and Systems Biotechnology
11 papers in training set
Top 0.2%
0.8%
28
Biochemistry
148 papers in training set
Top 2%
0.8%
29
eLife
5828 papers in training set
Top 66%
0.8%
30
Bioinformatics Advances
203 papers in training set
Top 5%
0.8%