Interpretable adenylation domain specificity prediction using protein language models
Adduri, A.; McNutt, A.; Ellington, C.; Suraparaju, K.; Fang, N.; Yan, D.; Krummenacher, B.; Li, S.; Bodden, C.; Xing, E.; Behsaz, B.; Koes, D. R.; Mohimani, H.
Show abstract
Natural products have long been a rich source of diverse and clinically effective drug candidates. Non-ribosomal peptides (NRPs), polyketides (PKs), and NRP-PK hybrids are three classes of natural products that display a broad range of bioactivities, including antibiotic, antifungal, anticancer, and immunosuppressant activities. However, discovering these compounds through traditional bioactivity-guided techniques is costly and time-consuming, often resulting in the rediscovery of known molecules. Consequently, genome mining has emerged as a high-throughput strategy to screen hundreds of thousands of microbial genomes to identify their potential to produce novel natural products. Adenylation domains play a key role in the biosynthesis of NRPs and NRP-PKs by recruiting substrates to incrementally build the final structure. We propose MASPR, a machine learning method that leverages protein language models for accurate and interpretable predictions of A-domain substrate specificities. MASPR demonstrates superior accuracy and generalization over existing methods and is capable of predicting substrates not present in its training data, or zero-shot classification. We use MASPR to develop Seq2Hybrid, an efficient algorithm to predict the structure of hybrid NRP-PK natural products from microbial genomes. Using Seq2Hybrid, we propose putative biosynthetic gene clusters for the orphan natural products Octaminomycin A, Dityromycin, SW-163B, and JBIR-39.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 94%
- A novel transformer-based platform for the prediction and design of biosynthetic gene clusters for (un)natural products 94%
- Deducing high-accuracy protein contact-maps from a triplet of coevolutionary matrices through deep residual convolutional networks 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.