Learning residue-level context for modeling protein-protein interactions
Zhang, Z.; Yang, Z.; Liu, A.; Yu, K.-H.; Zhao, J.; Yang, Y.; Neale, B.; Chen, S.
Show abstract
Protein language models (PLMs) enable prediction of protein properties by learning residue-level features from sequence, yet most PLM-based approaches to protein-protein interactions aggregate information across entire proteins, limiting resolution and interpretability. Here we present ReCLIP, a transformer-based framework that learns interaction-specific representations at the level of individual residues by combining intra-protein residue neighborhoods with residue-conditioned representations of interaction partners. We show that residue-centered context provides a general framework for modeling protein interactions across diverse biological settings. ReCLIP accurately predicts mutation-induced perturbations (AUROC = 0.973), generalizes to post-translational modifications that do not alter sequence (AUROC = 0.822), and enables zero-shot prediction of peptide-MHC binding across unseen alleles (AUROC up to 0.972). Analysis of learned residue neighborhoods reveals structurally and functionally coherent patterns aligned with known determinants of binding. Applied to clinically annotated genetic variants, ReCLIP identifies disease-associated interaction perturbations that link pathogenic variants to specific molecular interaction contexts. Our results establish a generalizable and interpretable framework for modeling protein interactions and provide insights into how residue-level context shapes interaction specificity and its perturbation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 97%
- Generalizable and scalable protein stability prediction with rewired protein generative models 97%
- Higher-order epistasis creates idiosyncrasy, confounding predictions in protein evolution 96%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- PSICHIC: physicochemical graph neural network for learning protein-ligand interaction fingerprints from sequence data 97%
- Predicting functional effect of missense variants using graph attention neural networks 96%
- Self-iterative multiple instance learning enables the prediction of CD4+ T cell immunogenic epitopes 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.