Back

SpY-C: Supervised Learning of Phosphopeptide Sequence Constraints Enables Global Prediction of SH2 Domain Binding

Kandoor, A.; Silva Oliveira, A. C.; Machida, K.; Blagoev, B.; Naegle, K. M.

2026-08-21 systems biology
10.64898/2026.08.18.744962 bioRxiv
Show abstract

Tyrosine kinase signaling for cell development and homeostasis in multicelluar organisms and a major biochemical contribution is by driving interactions between phosphorylated tyrosines (pY) and SH2 domain containing proteins. This assembly is so important to driving cell outcomes that a wide variety of experimental and computational approaches have been used to understand which SH2-pY interactions occur, which still remains a challenge given the immensity (more than 45,000 pY and 120 SH2 domains in the human proteome). Based on biophysical constraints suggested by comprehensive contact mapping, here, we ask whether an approach might consider first asking if pY sequences conform to the shared rules of SH2 domain recognition by developing a classification approach that combines diverse training data. A wide range of validation suggests this approach, SpY-C, can classify pY sites as having the potential, or not, to be involved in SH2 domain interactions. We find that a relatively small set of representative SH2 binders, integrated from different experimental techniques, provides good classification. We use this classifier to annotate the human phosphoproteome and individual experiments, to explore the consequences of using super-SH2 domain reagents for pY enrichment, and to analyze the effects of mutations in altering pY site function. SpY-C provides a helpful step to more rapidly annotating pY function and for possibly improving machine learning approaches focused on specific SH2-pY interactions downstream of a first pass classification approach.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.