Back

Mining of natural diversity enables efficient and expressible peptide asparaginyl ligases

Hemu, X.; Wenyu, D.; Qi, S.; Zhen, M.; Liao, M.; Zhao, Y.; Ma, J.; Hao, Y.; Jiang, H.; Lu, G.; Liew, C. W.; Chua, N.; Chen, H.; Hu, G.

2025-12-16 biochemistry
10.64898/2025.12.16.694575 bioRxiv
Show abstract

Peptide asparaginyl ligases (PALs) are powerful tools for protein engineering but are limited by natural rarity and poor expression. We mined 23 cyclotide-rich Viola species, uncovering 29 new PALs that expanded the known repertoire to 47. A dual-objective screen identified VdiPAL1 as the best-performed natural PAL, with twice efficiency of wt-VyPAL2 and 12 mg L-1 soluble expression in E. coli. A broad P2 specificity including Trp/Ile/Leu/Phe/Tyr/Met was discovered across diverse PALs, which enables sequential click-compatible liposome dual-functionalization. 1.8L[A] crystal structure of VdiPAL1 reveals a pre-organized near-attack conformation (NAC), supported by constant-pH MD simulations linking pH-dependent reactivity to NAC geometry. Our homology- and structure-based design yielded VyOpt1, a quintuple mutant of VyPAL2 with over 24-fold improved expression via enhanced cap-domain foldability in a single design-test cycle. This work expands the PAL family and demonstrates a transferable cap-domain-based engineering strategy, highlighting natural diversity as a powerful driver of enzyme discovery and optimization. TOC summaryMining of cyclotide-rich Viola genus expanded the total number of natural peptide asparaginyl ligases (PALs) to 47, including the discovery of highly expressible VdiPAL1. Its high-resolution structure provided new insights for mechanism, and its sequence guided us to a generalized PAL engineering strategy, leading to 24-fold increment in VyPAL2 expression.

Published in Nature Communications (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.