Prediction of protein-carbohydrate binding sites from protein primary sequence
Nawar, Q. F.; Nafi, M. M. I.; Islam, T. N.; Rahman, M. S.
Show abstract
BackgroundA protein is a large complex macromolecule that has a crucial role in performing most of the work in cells and tissues. It is made up of one or more long chains of amino acid residues. Another important biomolecule, after DNA and protein, is carbohydrate. Carbohydrates interact with proteins to run various biological processes. Several biochemical experiments exist to learn the protein-carbohydrate interactions, but they are expensive, time-consuming, and challenging. Therefore, developing computational techniques for effectively predicting protein-carbohydrate binding interactions from protein primary sequence has given rise to a prominent new field of research. ResultIn this study, we propose StackCBEmbed, an ensemble machine learning model to effectively classify protein-carbohydrate binding interactions at residue level. StackCBEmbed combines traditional sequence-based features along with features derived from a pre-trained transformer-based protein language model. To the best of our knowledge, ours is the first attempt to apply protein language model in predicting protein-carbohydrate binding interactions. StackCBEmbed achieved sensitivity and balanced accuracy scores of 0.730, 0.776 and 0.666, 0.742 in two separate independent test sets. This performance is superior compared to the earlier prediction models benchmarked in the same datasets. ConclusionWe thus hope that StackCBEmbed will discover novel protein-carbohydrate interactions and help advance the related fields of research. StackCBEmbed is freely available as Python scripts at https://github.com/nafiislam/StackCBEmbed.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 97%
- DBpred: A deep learning method for the prediction of DNA interacting residues in protein sequences 97%
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 96%
Similar papers in this journal
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 98%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 97%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 97%
Similar papers in this journal
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 96%
- Pred-AHCP: Robust feature selection enabled Sequence Specific Prediction of Anti-Hepatitis C Peptides via Machine Learning 96%
- Identification of Family-Specific Features in Cas9 and Cas12 Proteins: A Machine Learning Approach Using Complete Protein Feature Spectrum 96%
Similar papers in this journal
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 98%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 96%
- DrugForm-DTA: Towards real-world drug-target binding Affinity Model 96%
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 97%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 97%
- ToxinPred 3.0: An improved method for predicting the toxicity of peptides 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.