Back

CoMPHI: A Novel Composite Machine Learning Approach Utilizing Multiple FeatureRepresentation to Predict Hosts of Bacteriophages

Bodaka, S.; Malgonde, O.

2024-08-02 bioinformatics
10.1101/2024.07.29.604684 bioRxiv
Show abstract

Phage therapy has reemerged as a compelling alternative to antibiotics in treating bacterial infections, especially for superbugs that have developed antibiotic resistance. The challenge in the broader application of phage therapy is identifying host targets for the vast array of uncharacterized phages obtained through next-generation sequencing. To solve this issue, this paper introduces an innovative Composite Model for Phage Host Interaction, CoMPHI, to predict phage-host interactions by combining the accuracy of alignment-based methods with the efficiency and flexibility of machine learning techniques. The model initially generates multiple feature encodings from nucleotide and protein sequences of both phages and hosts to enhance prediction accuracies. It is further enriched by incorporating alignment scores between phage-phage, phage-host, and host-host, creating a composite model. During the 5-fold cross-validation, the composite model exhibited an Area Under the ROC Curve (AUC) of 94%, 96.4%, 96.5%, 96.6%, 96.6%, and 96.7% and accuracy of 92.3%, 93.3%, 93.6%, 94%, 94.9%, and 95.1% at the Species, Genus, Family, Order, Class, and Phylum levels, respectively. A comparative analysis revealed a 6-8% increase in model performance due to the inclusion of alignment scores. Additionally, an ablation study highlighted that including both nucleotide and protein sequences from both phages and hosts increased the prediction accuracy of the model. Another ablation study provided evidence that phage-host and host-host alignment scores, combined with phage-phage scores, equally contributed to enhancing the composite models performance. In conclusion, this paper presents a robust and comprehensive composite model advancing the use of phage therapy in modern medicine.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

1
Frontiers in Bioinformatics
49 papers in training set
Top 0.1%
10.6%
2
BMC Bioinformatics
457 papers in training set
Top 1.0%
7.8%
3
Bioinformatics
1204 papers in training set
Top 3%
6.7%
4
Bioinformatics Advances
203 papers in training set
Top 0.9%
5.5%
5
Computers in Biology and Medicine
128 papers in training set
Top 0.5%
5.5%
6
PeerJ
308 papers in training set
Top 1%
4.8%
7
GigaScience
212 papers in training set
Top 0.8%
4.3%
8
Frontiers in Microbiology
427 papers in training set
Top 2%
4.3%
9
Briefings in Bioinformatics
354 papers in training set
Top 2%
4.3%
50% of probability mass above
10
Frontiers in Genetics
230 papers in training set
Top 0.8%
4.0%
11
PLOS ONE
5266 papers in training set
Top 38%
3.2%
12
Journal of Theoretical Biology
162 papers in training set
Top 1%
2.4%
13
Scientific Reports
3612 papers in training set
Top 44%
2.4%
14
PLOS Computational Biology
1863 papers in training set
Top 13%
2.1%
15
Biology Methods and Protocols
61 papers in training set
Top 0.6%
2.1%
16
Computational Biology and Chemistry
28 papers in training set
Top 0.3%
1.9%
17
BioData Mining
22 papers in training set
Top 0.3%
1.7%
18
BMC Genomics
406 papers in training set
Top 5%
1.5%
19
NAR Genomics and Bioinformatics
242 papers in training set
Top 3%
1.4%
20
Computational and Structural Biotechnology Journal
242 papers in training set
Top 4%
1.4%
21
Journal of Computational Biology
48 papers in training set
Top 0.8%
1.1%
22
Viruses
332 papers in training set
Top 4%
1.1%
23
Biology
45 papers in training set
Top 0.4%
1.1%
24
Genomics
64 papers in training set
Top 1%
1.1%
25
International Journal of Molecular Sciences
494 papers in training set
Top 13%
1.0%
26
F1000Research
88 papers in training set
Top 5%
0.6%
27
Journal of Bioinformatics and Systems Biology
15 papers in training set
Top 0.4%
0.6%