Unsupervised protein language models learn patterns of enzyme function
Penner, M.; Lihan, M.; Bormke, H.; Nix, P.; Moscho, H.; Dupree, P.; Hollfelder, F.
Show abstract
While enormous amounts of sequence information have become available, assignment of sequence to a particular enzymatic function has remained elusive. Here we describe a framework that drives a general protein language model to find a target reaction without specific training, using an initial bridgehead protein. At the heart of this framework is PLM-clust, an algorithm that employs k-means on top of protein language model embeddings to convert sequence space into functional reservoirs of latent space, and samples from these clusters based on accelerated zero-shot scoring. We demonstrate PLM-clust in a recursive discovery process (with enzyme hit rates quickly rising to >90%), segmenting isofunctional reservoirs and exploring them in greater detail. This approach - exemplified for glycosyl hydrolases (a xylanase, >100-fold activity increase) and for imine reductases (IREDs, >100-fold increase in catalytic promiscuity profiles) - reliably brings about novel enzymes that are proficient at the catalytic task at hand, reaching deeply into sequence space with a majority of residues exchanged.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 95%
- PIFiA: Self-supervised Approach for Protein Functional Annotation from Single-Cell Imaging Data 95%
- AI-guided pipeline for protein-protein interaction drug discovery identifies a SARS-CoV-2 inhibitor 95%
Similar papers in this journal
- SHARK enables homology assessment in unalignable anddisordered sequences 97%
- Parametrically guided design of beta barrels and transmembrane nanopores using deep learning 96%
- Versatile NTP recognition and domain fusions expand the functional repertoire of the ParB-CTPase fold beyond chromosome segregation 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.