Back

Plastic-hydrolytic enzyme classification using explainable deep learning

Lee, W.-H.; Dumontet, L.; Jung, K.; Lee, H.; Thapa, G.; Oh, T.-J.; Kang, M.

2025-07-18 bioinformatics
10.1101/2025.07.14.664602 bioRxiv
Show abstract

The rapid accumulation of plastic waste has emerged as a critical environmental threat, driving the need for scalable and effective biodegradation solutions. Hydrolytic plastic-degrading enzymes (PDEs) offer a promising solution, yet their functional classification remains limited by insufficient annotations and enzymatic diversity. In this study, we present an explainable deep learning framework, PEPIC, to classify nine types of PDEs directly from protein sequences. Using a curated dataset of experimentally validated enzymes and an expanded homologous dataset, we built an explainable deep learning model based on convolutional neural networks (PEPIC) for plastic-degrading enzyme prediction. We benchmarked PEPICs performance against state-of-the-art approaches. First, PEPIC demonstrated statistically significant improvements in predictive performance compared to state-of-the-art methods. Second, PEPIC calculates contribution scores for each amino acid in the protein sequence, indicating their influence on the predictions. The model interpretation revealed that regions highlighted by high contribution scores matched conserved catalytic triads and substrate-binding clefts across PET-, PCL-, and PLA-degrading enzymes. Furthermore, structural modeling confirmed the trustworthiness of PEPICs predictions. Finally, PEPIC predicted an uncurated enzyme as a PET-degrading enzyme, which was biologically validated to hydrolyze bis(2-hydroxyethyl) terephthalate (BHET). These findings demonstrate that PEPIC provides accurate and trustworthy predictions of PDEs, facilitating the discovery of novel enzymes and supporting the development of sustainable plastic biodegradation technologies.

Matching journals

The top 12 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.