predPPI-GReMLIN: prediction of protein-protein interactions through mining of conserved bipartite graphs
Adeagbo, M. A.; Goncalves de Almeida, V. M.; Izidoro, S. C.; Silveira, S. d. A.
Show abstract
Protein-protein interactions (PPIs) play a central role in elucidating cellular mechanisms. However, a substantial gap remains in current prediction models, as they frequently overlook the structural and physicochemical context governing molecular binding, thereby limiting predictive accuracy. To address this limitation, we introduce predPPI-GReMLIN, a graph-based framework that represents protein-protein interfaces as bipartite graphs, integrating atomic-level physicochemical descriptors, spatial distance constraints, and conserved substructure mining to perform PPI prediction. The method employs a graph-search strategy to detect interface interaction patterns and conducts ligand swapping with complexes that share the same patterns as the query, enabling the prediction of novel interaction partners. Evaluations across multiple datasets--including CAMP (protein-peptide), Yeast (binary PPI), and TAGPPI (multi-class PPI)--demonstrate consistently strong predictive performance, achieving precision, recall, accuracy, and F1-scores exceeding 97% on binary classification benchmarks and surpassing state-of-the-art sequence- and structure-based approaches. Furthermore, incorporating solvent-accessible surface area (SASA)-derived features improved multi-class interaction-type classification accuracy to 57.74%. A case study on the SARS-CoV-2 spike-ACE2 complex further validated the approach: docking simulations using ligands predicted by our method reproduced native-like binding energetics comparable to redocking results. Additionally, a comparison between docking using a predPPI-GReMLIN-predicted ligand and docking using ligands selected at random from the PDB yielded a highly significant p-value (p = 5.1 x 10-9), indicating a robust statistical difference in binding performance. Collectively, these findings demonstrate predPPI-GReMLINs ability to capture conserved structural determinants of PPIs, providing a robust and interpretable framework for protein interaction prediction and ligand discovery. The dataset and source code for the experiments are publicly available at https://github.com/morufwork/predPPIGReMLIN.git.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DELPHI: accurate deep ensemble model for protein interaction sites prediction 98%
- Patch-DCA: Improved Protein Interface Prediction by utilizing Structural Information and Clustering DCA scores 98%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 97%
Similar papers in this journal
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 97%
- NRGSuite-Qt: A PyMOL plugin for high-throughput virtual screening, molecular docking, normal-mode analysis, the study of molecular interactions and the detection of binding-site similarities 97%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
Similar papers in this journal
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 98%
- Accurate prediction of residue-residue contacts across homo-oligomeric protein interfaces through deep leaning 97%
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 97%
Similar papers in this journal
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 97%
- Fast Local Alignment of Protein Pockets (FLAPP): A system-compiled program for large-scale binding site alignment 96%
- ArtiDock: accurate Machine Learning approach to protein-ligand docking optimized for high-throughput virtual screening 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.