SEHI-PPI: An End-to-End Sampling-Enhanced Human-Influenza Protein-Protein Interaction Prediction Framework with Double-View Learning
Yang, Q.; Fan, X.; Zhao, H.; Ma, Z.; Stanifer, M. L.; Bian, J.; Salemi, M.; Yin, R.
Show abstract
Influenza continues to pose significant global health threats, hijacking host cellular machinery through protein-protein interactions (PPIs), which are fundamental to viral entry, replication, immune evasion, and transmission. Yet, our understanding of these host-virus PPIs remains incomplete due to the vast diversity of viral proteins, their rapid mutation rates, and the limited availability of experimentally validated interaction data. Additionally, existing computational methods often struggle due to the limited availability of high-quality samples and their inability to effectively model the complex nature of host-virus interactions. To address these challenges, we present SEHI-PPI, an end-to-end framework for human-influenza PPI prediction. SEHI-PPI integrates a double-view deep learning architecture that captures both global and local sequence features, coupled with a novel adaptive negative sampling strategy to generate reliable and high-quality negative samples. Our method outperforms multiple benchmarks, including state-of-the-art large language models, achieving a superior performance in sensitivity (0.986) and AUROC (0.987). Notably, in a stringent test involving entirely unseen human and influenza protein families, SEHI-PPI maintains strong performance with an AUROC of 0.837. The model also demonstrates high generalizability across other human-virus PPI datasets, with an average sensitivity of 0.929 and AUROC of 0.928. Furthermore, AlphaFold3-guided case studies reveal that viral proteins predicted to target the same human protein cluster together structurally and functionally, underscoring the biological relevance of our predictions. These discoveries demonstrate the reliability of our SEHI-PPI framework in uncovering biologically meaningful host-virus interactions and potential therapeutic targets.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Learning interpretable cellular embedding for inferring biological mechanisms underlying single-cell transcriptomics 96%
- Scalable embedding fusion with protein language models: insights from benchmarking text-integrated representations 96%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 95%
Similar papers in this journal
- High-accuracy protein complex structure modeling based on sequence-derived structure complementarity 95%
- Integration of molecular coarse-grained model into geometric representation learning framework for protein-protein complex property prediction 95%
- Pretrainable Geometric Graph Neural Network for Antibody Affinity Maturation 95%
Similar papers in this journal
- Interpretable PROTAC degradation prediction with structure-informed deep ternary attention framework 97%
- Automatically Defining Protein Words for Diverse Functional Predictions Based on Attention Analysis of a Protein Language Model 95%
- Cross-modal Graph Contrastive Learning with Cellular Images 94%
Similar papers in this journal
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 95%
- A Multimodal Deep Learning Framework for Predicting PPI-Modulator Interactions 95%
- SSPSPredictor: A Sequence and Structure based Deep Learning Model for Predicting Phase-Separating Proteins 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.