PreDSLpmo V2.0: A deep learning-based prediction tool for functional annotation of lytic polysaccharide monooxygenases
Arulselvan, K. S.; Pulavendran, P.; Saravanan, V.; Yennamalli, R. M.
Show abstract
Lytic polysaccharide monooxygenases (LPMOs) are crucial enzymes that enhance the breakdown of polysaccharides, especially for biofuel production. Current computational tools for LPMO annotation are limited to a few families, leaving many unexplored. Here, we introduce PreDSLpmo v2.0, a deep learning-based tool that classifies and annotates the fast growing numbers of LPMOs across eight families using a curated dataset comprising over 30,000 LPMO sequences. To capture the compositional, physicochemical, and structural properties of these sequences, we extracted features using Python (iFeature) as well as R-based pipelines, generating over 13500 descriptors per sequence. Ensemble feature selection was used to identify significant features for binary and multiclass classification. To address data imbalance, we maintained a 1:1 ratio for all positive and negative sets during training and validation. A range of machine learning models were systematically trained and evaluated. An independent dataset was used to estimate the performance of the trained models. The multiclass Bi-LSTM model demonstrated the highest accuracy, robustness, and generalizability, outperforming feature-based approaches. We compared the performance of the model with Pfam, dbCAN3, and BlastP searches against the UniProtKB/swissprot database. The F1-score shows that the models predicted LPMO sequences are accurate. The reliability of predictions by Bi-LSTM model is comparable to that of dbCAN3 and often better, as confirmed by Pfam domain annotation and SignalP. The model, now deployed as a web server (https://predlpmo.in) for high-throughput, sequence-based functional annotation of LPMOs, provides a scalable and reliable solution for enzyme discovery in bioenergy and industrial biotechnology.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 95%
- TemStaPro: protein thermostability prediction using sequence representations from protein language models 95%
- RP3Net: a deep learning model for predicting recombinant protein production in Escherichia coli 95%
Similar papers in this journal
- ProtAlign-ARG: Antibiotic Resistance Gene Characterization Integrating Protein Language Models and Alignment-Based Scoring 96%
- Deep embeddings to comprehend and visualize microbiome protein space 95%
- PIPENN-EMB: ensemble net and protein embeddings generalise protein interface prediction beyond homology 95%
Similar papers in this journal
- From Signal to Symphony: Exploring 2D Sequence Representations for Protein Function Prediction 94%
- BERT-T6: Towards High-accuracy T6SS Bacterial Toxin Identification Using Protein Language Model 93%
- To Improve Protein Sequence Profile Prediction through Image Captioning on Pairwise Residue Distance Map 93%
Similar papers in this journal
- Protein language models can capture protein quaternary state 95%
- DeepSEA: an alignment-free deep learning tool for functional annotation of antimicrobial resistance proteins 94%
- CysPresso: A classification model utilizing deep learning protein representations to predict recombinant expression of cysteine-dense peptides 94%
Similar papers in this journal
- MoCETSE: A mixture-of-convolutional experts and transformer-based model for predicting Gram-negative bacterial secreted effectors 94%
- Positional SHAP (PoSHAP) for Interpretation of Machine Learning Models Trained from Biological Sequences 94%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.