Cataloging cysteines in ECOD domains using a protein language model
Yuan, R. D.; Durham, J.; Cong, Q.; Schaeffer, R. D. D.
Show abstract
Cysteine is among the most chemically versatile residues in the proteome, existing in three competing functional states: metal coordination, covalent disulfide bonding, and bioactive free thiols. Although these states can be readily assigned from experimentally determined protein structures using simple geometric criteria, accurately annotating them from predicted structures remains challenging. To bridge the gap between predicted structures and functional interpretation, we developed TriCyP (Tri-state Cysteine Predictor), an efficient two-layer neural network built on ESM-2 protein language model embeddings. On an independent benchmark set, TriCyP achieves near-perfect accuracy (AUROC = 0.99) and outperforms existing approaches for predicting both disulfide bonding and metal coordination. We applied TriCyP to classify 2.7 million cysteine residues across 0.9 million ECOD F70 representative domains. The resulting proteome-scale landscape recapitulates established biological patterns. Cysteines are enriched in eukaryotes: disulfide-bonded states are concentrated in extracellular proteins, and metal-coordinating cysteines peak in nuclear proteins owing to the abundance of zinc-finger transcription factors. We further demonstrate the utility of cysteine-state annotation through two pilot studies. First, predicted disulfide-forming cysteines lacking a corresponding structural partner in AlphaFold models may identify either regions of elevated structural uncertainty or latent inter-protein disulfide bonds that stabilize protein-protein interactions. Second, systematic analysis of known and predicted metal-coordinating cysteines across ECOD homologous groups uncovers previously unrecognized metal-binding protein families. This proteome-wide catalog of cysteine states is available as a community resource (http://prodata.swmed.edu/tricyp) and will be integrated into future ECOD releases.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identification of Iron-Sulfur (Fe-S) and Zn-binding Sites Within Proteomes Predicted by DeepMind's AlphaFold2 Program Dramatically Expands the Metalloproteome 96%
- The 3D modules of enzyme catalysis: deconstructing active sites into distinct functional entities 93%
- Assembly of Protein Complexes In and On the Membrane with Predicted Spatial Arrangement Constraints 93%
Similar papers in this journal
- Deep Local Analysis deconstructs protein-protein interfaces and accurately estimates binding affinity changes upon mutation 96%
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 96%
- Mapping the space of protein binding sites with sequence-based protein language models 95%
Similar papers in this journal
- Pervasive, conserved secondary structure in highly charged protein regions 95%
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 95%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.