PhageTransformer - scalable and accurate host assignments for bacteriophages
Siemers, M.; Lopez, J. L.; Dutilh, B. E.
Show abstract
Bacteriophages can only be understood through their interactions with bacterial hosts. As environmental sequencing efforts expanded, the number of available phage genome sequences has exploded, yet the vast majority of these sequences lack host information. Predicting the host of a newly observed phage is therefore a key challenge in virology. Several computational tools can predict phage-host relationships from genomic data, but they share notable limitations: (1) the number of different hosts that can be predicted remains relatively restricted; (2) tools tend to assign confident host predictions to non-viral input sequences; and (3) most tools have a trade-off between accuracy and speed. Here we present PhageTransformer (PT), a deep learning model for phage-host prediction that addresses these limitations. We benchmark PT against existing tools on 3,881 independent phage-host pairs from GenBank and public HiC data, and demonstrate that it achieves competitive or superior prediction accuracy at greatly reduced runtime.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Virus-Host Interactions Predictor (VHIP): machine learning approach to resolve microbial virus-host interaction networks 93%
- A deep learning approach to real-time HIV outbreak detection using genetic data 92%
- MoCETSE: A mixture-of-convolutional experts and transformer-based model for predicting Gram-negative bacterial secreted effectors 92%
Similar papers in this journal
- nf-core/viralmetagenome: A Novel Pipeline for Untargeted Viral Genome Reconstruction 94%
- PHIST: fast and accurate prediction of prokaryotichosts from metagenomic viral sequences 93%
- PHIStruct: Improving phage-host interaction prediction at low sequence similarity settings using structure-aware protein embeddings 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.