Extending the application of the SCA/Sectors method for the identification of domain boundaries and subtype specific residues in multi-domain biosynthetic proteins: Application to Polyketide Synthases
Oruc, T.; Thomas, C. M.; Winn, P. J.
Show abstract
Polypeptides with multiple enzyme domains, such as type I polyketide synthases, produce chemically complex compounds that are difficult to produce via conventional chemical synthesis and are often pharmaceutically or otherwise commercially valuable. Engineering polyketide synthases, via domain swapping and/or site directed mutagenesis, in order to generate novel polyketides, has tended to produce either low yields of product or no product at all. The success of such experiments may be limited by our inability to predict the key functional residues and boundaries of protein domains. Computational tools to identify the boundaries and the residues determining the substrate specificity of domains could reduce the trial and error involved in engineering multi-domain proteins. In this study we use statistical coupling analysis to identify networks of co-evolving residues in type I polyketide synthases, thereby predicting domain boundaries. We extend the method to predicting key residues for enzyme substrate specificity. We introduce bootstrapping calculations to test the relationship between sequence length and the number of sequences needed for a robust analysis. Our results show no simple predictor of the number of sequences needed for an analysis, which can be as few as a hundred and as many as a few thousand. We find that polyketide synthases contain multiple networks of co-substituting residues: some are intradomain but most multiple domains. Some networks of coupled residues correlate with specific functions such as the substrate specificity of the acyl transferase domain, the stereo chemistry of the ketoreductase domain, or domain boundaries that are consistent with experimental data. Our extension of the method provides a ranking of the likely importance of these residues to enzyme substrate specificity, allowing us to propose residues for further mutagenesis work. We conclude that analysis of co-evolving networks of residues is likely to be an important tool for re-engineering multi-domain proteins. Author summaryMany important compounds such as antibiotics or food flavourings are produced naturally by molecular factories within plant, fungal and bacterial cells. These molecular factories typically comprise a complex of multiple interacting enzymes, each enzyme being a stage in a molecular production line. Often the enzymes are connected together as subsections of the same amino acid chain, i.e. protein, with the amino acid chain folding into the separate functional enzymatic domains that comprise the production line. Polyketide synthases are such multi-domain proteins, and their products often have antibacterial, antifungal and antitumoric effects. Engineering polyketide synthases thus has the potential to produce novel drug candidates. We applied and developed statistical approaches to detect where in an amino acid sequence the boundaries are between different domains, potentially allowing these regions to be swapped around for the synthesis of novel compounds. We used the same approaches to identify parts of the amino acid chain important for the function of different types of domain, pointing to how they might be modified to make novel compounds. These analyses agree with published experimental data and allow us to make novel predictions, which we expect to help experimentalists produce novel compounds of commercial and pharmaceutical interest.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dynamic coupling of residues within proteins as a mechanistic foundation of many enigmatic pathogenic missense variants 95%
- Integrating structure-based machine learning and co-evolution to investigate specificity in plant sesquiterpene synthases 94%
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 94%
Similar papers in this journal
- Larger active site in an ancestral hydroxynitrile lyase increases catalytically promiscuous esterase activity 95%
- Crystal structure of β-L-arabinobiosidase belonging to glycoside hydrolase family 121 94%
- Conserved intramolecular networks in GDAP1 are closely connected to CMT-linked mutations and protein stability 93%
Similar papers in this journal
- A computational study of the fold and stability of cytochrome c with implications for disease 94%
- Diversity, structure-function relationships and evolution of cell wall-binding domains of staphylococcal phage endolysins 93%
- A Mathematical Genomics Perspective on the Moonlighting Role of Glyceraldehyde-3-Phosphate Dehydrogenase (GAPDH) 93%
Similar papers in this journal
- Evolutionary progression of collective mutations in Omicron sub-lineages towards efficient RBD-hACE2: allosteric communications between and within viral and human proteins 94%
- Structural Insight into the Function of Human Peptidyl Arginine Deiminase 6 93%
- The Rad52 superfamily as seen by AlphaFold 93%
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 92%
- L-shape distribution of the relative substitution rate (c micro) observed for SARS-COV-2 genome, inconsistent with the selectionist theory, the neutral theory and the nearly neutral theory but a near-neutral balanced selection theory: implication on neutralist-selectionist debate 91%
- Uncovering Co-regulatory Modules and Gene Regulatory Networks in the Heart through Machine Learning-based Analysis of Large-scale Epigenomic Data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.