Back

Phylogenetic tree-aware positive-unlabeled deep metric learning for phage-host interaction identification

Zhang, Y.-z.; Tobias, B.; Imoto, S.

2025-12-31 bioinformatics
10.64898/2025.12.31.696981 bioRxiv
Show abstract

Phages are viruses that infect bacteria and play essential roles in shaping microbial communities. Identifying phage-host interactions (PHIs) is crucial for understanding infection dynamics and developing phage-based therapeutic strategies. Recent deep learning approaches have shown great promise for PHI prediction; however, their performance remains constrained by the limited number of experimentally validated positive pairs and the overwhelming abundance of unlabeled or non-validated samples. Moreover, most existing models overlook higher-level phylogenetic relationships among hosts, which could provide valuable structural priors for guiding representation learning. To address these challenges, we propose a phylogenetic tree-aware positive-unlabeled deep metric learning framework for phage-host interaction (PHI) identification. Unlike traditional approaches that train classification models to strictly separate positive and negative phage-host pairs, the proposed method learns representations under supervision from both confirmed positive PHIs and host phylogenetic tree constraints on non-positive samples. The proposed method can seamlessly formalize contrastive learning and deep metric learning within the same framework that explicitly optimizes PHI encoders with biological constraints in the learning functions. We show that this metric learning formulation outperforms conventional contrastive learning approaches that enforce separation between positive and negative samples without consistently aligning the learned representations with evolutionary distances. Experiments on the Cherry benchmark dataset and metagenome Hi-C multi-host dataset demonstrate that our approach enhances species-level prediction accuracy, improves cross-host generalization, and yields more interpretable representations of phage-host relationships.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.