Back

PreDSLpmo V2.0: A deep learning-based prediction tool for functional annotation of lytic polysaccharide monooxygenases

Arulselvan, K. S.; Pulavendran, P.; Saravanan, V.; Yennamalli, R. M.

2025-06-07 bioinformatics
10.1101/2025.06.05.658061 bioRxiv
Show abstract

Lytic polysaccharide monooxygenases (LPMOs) are crucial enzymes that enhance the breakdown of polysaccharides, especially for biofuel production. Current computational tools for LPMO annotation are limited to a few families, leaving many unexplored. Here, we introduce PreDSLpmo v2.0, a deep learning-based tool that classifies and annotates the fast growing numbers of LPMOs across eight families using a curated dataset comprising over 30,000 LPMO sequences. To capture the compositional, physicochemical, and structural properties of these sequences, we extracted features using Python (iFeature) as well as R-based pipelines, generating over 13500 descriptors per sequence. Ensemble feature selection was used to identify significant features for binary and multiclass classification. To address data imbalance, we maintained a 1:1 ratio for all positive and negative sets during training and validation. A range of machine learning models were systematically trained and evaluated. An independent dataset was used to estimate the performance of the trained models. The multiclass Bi-LSTM model demonstrated the highest accuracy, robustness, and generalizability, outperforming feature-based approaches. We compared the performance of the model with Pfam, dbCAN3, and BlastP searches against the UniProtKB/swissprot database. The F1-score shows that the models predicted LPMO sequences are accurate. The reliability of predictions by Bi-LSTM model is comparable to that of dbCAN3 and often better, as confirmed by Pfam domain annotation and SignalP. The model, now deployed as a web server (https://predlpmo.in) for high-throughput, sequence-based functional annotation of LPMOs, provides a scalable and reliable solution for enzyme discovery in bioenergy and industrial biotechnology.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.