Optimal feature selection and software tool development for bacteriocin prediction
Akhter, S.; Miller, J.
Show abstract
Antibiotic resistance is a major public health concern around the globe. As a result, researchers always look for new compounds to develop new antibiotic drugs for combating antibiotic-resistant bacteria. Bacteriocin becomes a promising antimicrobial agent to fight against antibiotic resistance, due to its narrow killing spectrum. Sequence matching methods are widely used to identify bacteriocins by comparing them with the known bacteriocin sequences; however, these methods often fail to detect new bacteriocin sequences due to sequences high diversity. The ability to use a machine learning approach can help find new highly dissimilar bacteriocins for developing highly effective antibiotic drugs. The aim of this work is to identify optimal sets of features and develop a machine learning-based software tool for predicting bacteriocin protein sequences with high accuracy. We extracted potential features from known bacteriocin and non-bacteriocin sequences by considering the physicochemical and structural properties of the protein sequences. Then we reduced the feature set using statistical justifications and recursive feature elimination technique. Finally, we built support vector machine (SVM) and random forest (RF) models using the selected features and our models can achieve accuracy up to 95.54%. We compared the performance of our method with a popular sequence matching-based approach and a deep learning-based method. We also developed a software tool called Bacteriocin Prediction (BacPred) that implements the prediction model using the optimal set of features obtained from this study. The software package and its user manual are available at https://github.com/suraiya14/ML_bacteriocins/BacPred.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 96%
- Rprot-Vec: A deep learning approach for fast protein structure similarity calculation 95%
- PDAUG - a Galaxy based toolset for peptide library analysis, visualization, and machine learning modeling 95%
Similar papers in this journal
- MACI: A machine learning-based approach to identify drug classes of antibiotic resistance genes from metagenomic data 97%
- ToxinPred 3.0: An improved method for predicting the toxicity of peptides 96%
- A method for predicting linear and conformational B-cell epitopes in an antigen from its primary sequence 96%
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 96%
- Antivirals for Monkeypox Virus: Proposing an Effective Machine/Deep Learning Framework 95%
- Heterologous expression and characterization of mutant cellulase from indigenous strain of Aspergillus niger 94%
Similar papers in this journal
- A Novel Riboswitch Classification based on Imbalanced Sequences achieved by Machine Learning 95%
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 94%
- DeepHE: Accurately Predicting Human Essential Genes based on Deep Learning 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.