SYNTAX Reads the Motif and Domain Grammar of the Androgen Receptor Proximal Interactome
Eng, J. K.; Wright, M. E.
Show abstract
A receptors interactome is usually reported as a protein list, yet the sequence rules that organize it have never been read out. We introduce SYNTAX, a scan of 356 short linear motif classes and 573 protein domains that corrects the length, burden, and degeneracy artifacts of motif counting, and apply it to 4,751 androgen receptor (AR)-proximal interacting proteins (AR-PIPs) across three subcellular compartments and an androgen time course. ARs signature coactivator motif, LXXLL, is the most depleted class, while the neighborhood speaks a disorder-resident signaling vocabulary of SUMOylation sites, SPOP and SIAH degrons, and nuclear-import signals. The paradox resolves once incidental motifs are removed. ARs coactivator core forms a small, 9-fold-enriched inner shell within a signaling periphery. The grammar reproduces in an estrogen receptor-beta interactome (r = 0.89) and generalizes to the 5-HT2A serotonin receptor and EGFR. SYNTAX interfaces with the Predictive Proximal Proteome (PPP) Database, a queryable AR-PIP resource.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 92%
- Inferring high-dimensional pathways of trait acquisition in evolution and disease 92%
- Geometric Sketching Compactly Summarizes the Single-Cell Transcriptomic Landscape 91%
Similar papers in this journal
Similar papers in this journal
- Proteome-level assessment of origin, prevalence and function of Leucine-Aspartic Acid (LD) motifs 94%
- SHEPHARD: a modular and extensible software architecture for analyzing and annotating large protein datasets 93%
- BindPred: A Framework for Predicting Protein-Protein Binding Affinity from Language Model Embeddings 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.