Discovery of a phenazine thiol conjugase from sparse data using genome-informed machine learning
Newman, D. K.; Shan, X.; Trindade, I. B.; Glasser, N. R.; Thalhammer, K. O.; Scurria, M.; Mora, A.; Conway, S. J.
Show abstract
Machine learning has enabled powerful biological discoveries using models trained on large datasets. However, for many important biological questions, such as identifying enzymes that transform understudied substrates, sparsity of training data is often a major bottleneck. Here, using phenazine natural products as a case study, we show that integrating genome-informed data augmentation with contrastive learning in protein language space enables identification of phenazine-interacting proteins starting from only 14 known phenazine modifying sequences. Applying this framework led to the discovery of PTC (Phenazine-Thiol Conjugase), the first enzyme known to catalyze phenazine thioconjugation, a phenazine modification reaction long observed but previously presumed to occur only through non-enzymatic chemistry. In silico simulation and experimental measurements demonstrate that PTC binds to both phenazine and glutathione as substrates. Recombinant expression and biochemical characterization reveal that PTC promotes glutathione-dependent modification of phenazines, yielding distinct reaction outcomes that depend on substrate identity. Although thiol-conjugated phenazine products exhibit reduced toxicity to bacterial cells, deletion of the gene encoding PTC does not confer a strong fitness disadvantage, illustrating how direct learning of sequences can uncover relevant enzymes that might evade phenotype-based genetic screens. Together, these results demonstrate that coupling comparative genomics with protein machine learning can convert "small data" typically outside the scope of machine learning into actionable predictive power, thereby facilitating enzyme discovery. SignificanceMachine learning excels when large, well-labeled datasets are available, yet many biologically important problems lack sufficient experimental data to support such approaches to discovery. This limitation is particularly acute for identifying enzymes acting on rare or understudied substrates. Here, we show that genomic organization can be leveraged as an additional source of biological information to address data sparsity. Starting with only 14 enzymes experimentally shown to modify phenazines, we developed a model identifying phenazine-interacting enzymes by integrating genome-informed data augmentation with protein machine learning. Guided by the model, we discovered the first enzyme known to catalyze thioconjugation modifications of phenazines, demonstrating a simple yet powerful strategy for extracting predictive insight from sparse biological knowledge.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enzymatic carbon-fluorine bond cleavage by human gut microbes 97%
- Structural basis for divergent and convergent evolution of catalytic machineries in plant aromatic amino acid decarboxylase proteins 96%
- Structural Basis for Iterative Methylation by a Cobalamin-dependent Radical S-Adenosylmethionine Enzyme in Cystobactamids Biosynthesis 96%
Similar papers in this journal
Similar papers in this journal
- The S-lignin O-demethylase SyoA: Structural insights into a new class of heme peroxygenase enzymes 96%
- Functional control of a 0.5 MDa TET aminopeptidase by a flexible loop revealed by MAS NMR 95%
- Non-consecutive enzyme interactions within TCA cycle supramolecular assembly regulate carbon-nitrogen metabolism 95%
Similar papers in this journal
- Biosynthesis of novel desferrioxamine derivatives requires unprecedented crosstalk between separate NRPS-independent siderophore pathways 95%
- Critical analysis of polycyclic tetramate macrolactam biosynthetic cluster phylogeny and functional diversity 94%
- Bacterial-like nonribosomal peptide synthetases produce cyclopeptides in the zygomycetous fungus Mortierella alpina 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.