Back

Data-Efficient Exploration of Enzyme Function Using Family-Specific Machine Learning

Ahmed, F. H.; Bender, A.; Wijesinghe, A.; Zhu, A.; Zhang, L.; Gebbie, L.; Marsh, A.; Ishitate, C.; Holdsworth, W.; Jones, C.; Warden, A. C.; Power, H.; Ong, C. S.; Steinberg, D. M.; Speight, R. E.

2026-06-03 bioengineering
10.64898/2026.06.02.729712 bioRxiv
Show abstract

Enzymes are essential biocatalysts across diverse industries, driving demand for high-performing variants. Foundation models are attractive for guiding enzyme discovery, but often lack the resolution to model subtle variations driving function within homologous families. Navigating these rugged functional landscapes to identify elite variants remains challenging and experimentally costly, even when guided by such models. Here we show that coupling dense, family-specific experimental screening with targeted, sequence-based deep learning provides a data-efficient discovery strategy. We experimentally screened 1,513 natural homologues from an esterase superfamily (>7,500 assays) and used this functional landscape to train task-specific models that predict activity, thermostability, and substrate specificity from sequence alone. Prospective experimental validation of previously untested sequences demonstrated that these task-specific models significantly outperformed generalist pre-trained and physics-based models in enriching for target traits. Residue-level attribution further indicated that the models captured sequence patterns consistent with underlying structural features. Finally, retrospective simulations showed that iterative retraining compresses the search space, discovering 60% of top-tier hits using nearly half the samples required by pre-trained baseline models. Together, these results highlight that machine learning can provide mechanistic insight, and that integrating targeted data acquisition with iterative machine learning provides a more data-efficient discovery strategy than relying on generic model scale.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

1
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 2%
12.7%
2
Nature Communications
5641 papers in training set
Top 14%
12.7%
3
Nature Chemistry
42 papers in training set
Top 0.1%
9.6%
4
ACS Catalysis
18 papers in training set
Top 0.1%
6.7%
5
Cell Systems
201 papers in training set
Top 0.5%
6.7%
6
PLOS Computational Biology
1863 papers in training set
Top 8%
4.3%
50% of probability mass above
7
Nature Machine Intelligence
70 papers in training set
Top 0.7%
4.0%
8
eLife
5828 papers in training set
Top 32%
3.4%
9
Angewandte Chemie International Edition
93 papers in training set
Top 0.6%
3.2%
10
Science Advances
1243 papers in training set
Top 11%
3.2%
11
Journal of Chemical Information and Modeling
238 papers in training set
Top 1%
2.4%
12
ACS Central Science
71 papers in training set
Top 0.5%
2.4%
13
Nature Chemical Biology
119 papers in training set
Top 1%
2.4%
14
Nature
645 papers in training set
Top 6%
2.1%
15
Advanced Science
286 papers in training set
Top 4%
1.9%
16
Journal of the American Chemical Society
217 papers in training set
Top 2%
1.7%
17
Science
477 papers in training set
Top 6%
1.5%
18
Communications Chemistry
48 papers in training set
Top 1%
1.1%
19
Computational and Structural Biotechnology Journal
242 papers in training set
Top 5%
1.1%
20
Communications Biology
993 papers in training set
Top 26%
1.0%
21
Nucleic Acids Research
1281 papers in training set
Top 13%
1.0%
22
Neuron
337 papers in training set
Top 5%
1.0%
23
Trends in Biotechnology
12 papers in training set
Top 0.2%
0.9%
24
Nature Methods
385 papers in training set
Top 6%
0.9%
25
Protein Science
246 papers in training set
Top 4%
0.8%
26
Cell Chemical Biology
94 papers in training set
Top 2%
0.8%
27
Cell Reports
1498 papers in training set
Top 27%
0.8%
28
Chemical Science
73 papers in training set
Top 2%
0.6%