Back

predPPI-GReMLIN: prediction of protein-protein interactions through mining of conserved bipartite graphs

Adeagbo, M. A.; Goncalves de Almeida, V. M.; Izidoro, S. C.; Silveira, S. d. A.

2025-12-06 bioinformatics
10.64898/2025.12.03.692055 bioRxiv
Show abstract

Protein-protein interactions (PPIs) play a central role in elucidating cellular mechanisms. However, a substantial gap remains in current prediction models, as they frequently overlook the structural and physicochemical context governing molecular binding, thereby limiting predictive accuracy. To address this limitation, we introduce predPPI-GReMLIN, a graph-based framework that represents protein-protein interfaces as bipartite graphs, integrating atomic-level physicochemical descriptors, spatial distance constraints, and conserved substructure mining to perform PPI prediction. The method employs a graph-search strategy to detect interface interaction patterns and conducts ligand swapping with complexes that share the same patterns as the query, enabling the prediction of novel interaction partners. Evaluations across multiple datasets--including CAMP (protein-peptide), Yeast (binary PPI), and TAGPPI (multi-class PPI)--demonstrate consistently strong predictive performance, achieving precision, recall, accuracy, and F1-scores exceeding 97% on binary classification benchmarks and surpassing state-of-the-art sequence- and structure-based approaches. Furthermore, incorporating solvent-accessible surface area (SASA)-derived features improved multi-class interaction-type classification accuracy to 57.74%. A case study on the SARS-CoV-2 spike-ACE2 complex further validated the approach: docking simulations using ligands predicted by our method reproduced native-like binding energetics comparable to redocking results. Additionally, a comparison between docking using a predPPI-GReMLIN-predicted ligand and docking using ligands selected at random from the PDB yielded a highly significant p-value (p = 5.1 x 10-9), indicating a robust statistical difference in binding performance. Collectively, these findings demonstrate predPPI-GReMLINs ability to capture conserved structural determinants of PPIs, providing a robust and interpretable framework for protein interaction prediction and ligand discovery. The dataset and source code for the experiments are publicly available at https://github.com/morufwork/predPPIGReMLIN.git.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.