Protein language models accelerate the discovery of Plastic-Degrading Enzymes
Medina-Ortiz, D.; Alvarez-Saravia, D.; Soto-Garcia, N.; Sandoval-Vargas, D.; Aldridge, J.; Rodriguez, S.; Andrews, B.; Asenjo, J. A.; Daza, A.
Show abstract
Plastic pollution presents a critical environmental challenge, necessitating innovative and sustainable solutions. In this context, biodegradation using microorganisms and enzymes offers an environmentally friendly alternative. This work introduces an AI-driven frame-work that integrates machine learning (ML) and generative models to accelerate the discovery and design of plastic-degrading enzymes. By leveraging pre-trained protein language models and curated datasets, we developed seven ML-based binary classification models to identify enzymes targeting specific plastic substrates, achieving an average accuracy of 89%. The framework was applied to over 6,000 enzyme sequences from the RemeDB to classify enzymes targeting diverse plastics, including PET, PLA, and Nylon. Besides, generative learning strategies combined with trained classification models in this work were applied for de novo generation of PET-degrading enzymes. Structural bioinformatics validated potential candidates through in-silico analysis, highlighting differences in physicochemical properties between generated and experimentally validated enzymes. Moreover, generated sequences exhibited lower molecular weights and higher aliphatic indices, features that may enhance interactions with hydrophobic plastic substrates. These findings highlight the utility of AI-based approaches in enzyme discovery, providing a scalable and efficient tool for addressing plastic pollution. Future work will focus on experimental validation of promising candidates and further refinement of generative strategies to optimize enzymatic performance.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Enzymatic degradation of biofilm by metalloprotease from Microbacterium sp. SKS10 92%
- A computational framework to identify metabolic engineering strategies for the co-production of metabolites 91%
- Pseudomonas mRNA 2.0: Boosting Gene Expression Through Enhanced mRNA Stability and Translational Efficiency 90%
Similar papers in this journal
- Microbial Metabolic Enzymes, Pathways and Microbial Hosts for Co-Metabolic Degradation of Organic Micropollutants in Wastewater 93%
- L-norepinephrine Induces Community Shift, Oxidative Stress Response, Metabolic Reprogramming, and Virulence Potential in Wastewater Microbiomes 93%
- Substrate promiscuity of xenobiotic-transforming hydrolases from stream biofilms impacted by treated wastewater 92%
Similar papers in this journal
- Low cost and sustainable hyaluronic acid production in a manufacturing platform based on Bacillus subtilis 3NA strain 92%
- The biochemical properties of a novel paraoxonase-like enzyme in Trichoderma atroviride strain T23 involved in the degradation of 2,2-dichlorovinyl dimethyl phosphate 92%
- Production of nonulosonic acids in the extracellular polymeric substances of Candidatus Accumulibacter phosphatis 91%
Similar papers in this journal
- Probing Specificities of Alcohol Acyltransferases for Designer Ester Biosynthesis with a High-Throughput Microbial Screening Platform 94%
- A generalized machine-learning aided method for targeted identification of industrial enzymes from metagenome: a xylanase temperature dependence case study 92%
- Microfluidic single-cell scale-down bioreactors: A proof-of-concept for the growth of Corynebacterium glutamicum at oscillating pH values 91%
Similar papers in this journal
- Machine learning-assisted medium optimization revealed the discriminated strategies for improved production of the foreign and native metabolites 90%
- Implementation of a Clostridium luticellarii genome-scale model for upgrading syngas fermentations 90%
- FAMeDB: A curated Database for the analysis of Fungal Aromatic Compound Metabolism 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.