Back

Unsupervised protein language models learn patterns of enzyme function

Penner, M.; Lihan, M.; Bormke, H.; Nix, P.; Moscho, H.; Dupree, P.; Hollfelder, F.

2026-04-23 synthetic biology
10.64898/2026.04.23.720319 bioRxiv
Show abstract

While enormous amounts of sequence information have become available, assignment of sequence to a particular enzymatic function has remained elusive. Here we describe a framework that drives a general protein language model to find a target reaction without specific training, using an initial bridgehead protein. At the heart of this framework is PLM-clust, an algorithm that employs k-means on top of protein language model embeddings to convert sequence space into functional reservoirs of latent space, and samples from these clusters based on accelerated zero-shot scoring. We demonstrate PLM-clust in a recursive discovery process (with enzyme hit rates quickly rising to >90%), segmenting isofunctional reservoirs and exploring them in greater detail. This approach - exemplified for glycosyl hydrolases (a xylanase, >100-fold activity increase) and for imine reductases (IREDs, >100-fold increase in catalytic promiscuity profiles) - reliably brings about novel enzymes that are proficient at the catalytic task at hand, reaching deeply into sequence space with a majority of residues exchanged.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.