Prediction of inhibitory peptides against E. coli with desired MIC value
Bajiya, N.; Kumar, N.; Raghava, G. P. S.
Show abstract
In the past, several methods have been developed for predicting antibacterial and antimicrobial peptides, but only limited attempts have been made to predict their minimum inhibitory concentration (MIC) values. In this study, we trained our models on 3,143 peptides and validated them on 786 peptides whose MIC values have been determined experimentally against Escherichia coli (E. coli). The correlational analysis reveals that the Composition Enhanced Transition and Distribution (CeTD) attributes strongly correlate with MIC values. We initially employed the similarity search strategy utilizing BLAST to estimate MIC values of peptides but found it inadequate for prediction. Next, we developed machine learning techniques-based regression models using a wide range of features, including peptide composition, binary profile, and embeddings of large language models. We implemented feature selection techniques like minimum Redundancy Maximum Relevance (mRMR) to select the best relevant features for developing prediction models. Our Random forest-based regressor, based on selected features, achieved a correlation coefficient (R) of 0.78, R-squared (R{superscript 2}) of 0.59, and a root mean squared error (RMSE) of 0.53 on the validation dataset. Our best model outperforms the existing methods when benchmarked on an independent dataset of 498 inhibitory peptides of E. coli. One of the major features of the web-based platform EIPpred developed in this study is that it allows users to identify or design peptides that can inhibit E. coli with the desired MIC value (https://webs.iiitd.edu.in/raghava/eippred). HighlightsO_LIPrediction of MIC value of peptides against E.coli. C_LIO_LIAn independent dataset was generated for comparison. C_LIO_LIFeature selection using the mRMR method. C_LIO_LIA regressor method for designing novel inhibitory peptides. C_LIO_LIA web server and standalone package for predicting the inhibitory activity of peptides. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- ToxinPred 3.0: An improved method for predicting the toxicity of peptides 98%
- Prediction and scanning of IL-5 inducing peptides using alignment-free and alignment-based method 97%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 96%
Similar papers in this journal
- Pred-AHCP: Robust feature selection enabled Sequence Specific Prediction of Anti-Hepatitis C Peptides via Machine Learning 98%
- Predicting Antimicrobial Activity for Untested Peptide-Based Drugs Using Collaborative Filtering and Link Prediction 96%
- WITHDRAWN: The Use of DeepQSAR Models for The Discovery of Peptides with Enhanced Antimicrobial and Antibiofilm Potential. 95%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 97%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 95%
- DrugForm-DTA: Towards real-world drug-target binding Affinity Model 95%
Similar papers in this journal
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 97%
- PDAUG - a Galaxy based toolset for peptide library analysis, visualization, and machine learning modeling 96%
- CysPresso: A classification model utilizing deep learning protein representations to predict recombinant expression of cysteine-dense peptides 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.