Back

SPPIDER-seq: Sequence-based partner-aware predictor of protein-protein interaction sites

Porollo, A.; Jadhav, O.; Alvarez, A.; Chen, J.

2026-04-30 bioinformatics
10.64898/2026.04.28.721359 bioRxiv
Show abstract

MotivationSequence-based protein-protein interaction (PPI) site predictors typically analyze proteins in isolation, neglecting partner-specific context that is critical for interface specificity, particularly in transient and disordered interactions. ResultsWe introduce SPPIDER-seq, a partner-aware PPI site prediction framework that combines pretrained ESM-2 embeddings with a cross-attention architecture to enable residue-level conditioning on interacting partners. We curated non-redundant protein-peptide interaction datasets from BioLiP and used them to train and benchmark two complementary models: a receptor-centric model optimized for structured interfaces and a peptide-centric model tailored to disordered, motif-driven binding. On blind benchmarks, SPPIDER-seq achieved AUROC values up to 0.797 and MCC values up to 0.269, outperforming AlphaFold3 on peptide-mediated and disordered interfaces while remaining complementary on globular complexes. Application to 341 TP53 interaction partners revealed coherent, partner-specific interface patterns across both structured and intrinsically disordered regions. Availability and ImplementationSPPIDER-seq models, datasets, and the Python code are freely available on the web at: https://github.com/aporollo-lab/SPPIDER-seq and archived on Zenodo at DOI: 10.5281/zenodo.19835990, corresponding to GitHub release v2.0-manuscript. ContactDr. Aleksey Porollo - porollay@ucmail.uc.edu Dr. Jichao Chen - jichao.chen@cchmc.org Supplementary InformationAvailable online with the manuscript.

Published in Bioinformatics (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.