Disentangling the CHAOS of intrinsic disorder in human proteins
de Vries, I.; Bak, J.; Alvarez Salmoral, D.; Xie, R.; Borza, R.; Konijnenberg, M.; Perrakis, A.
Show abstract
Most proteins consist of both folded domains and Intrinsically Disordered Regions (IDRs). However, the widespread occurrence of intrinsic disorder in human proteins, along with its characteristics, is often overlooked by the broader communities of structural and molecular biologists. Building on the MobiDB database of intrinsic disorder in proteins, here we develop a comprehensive dataset (Comprehensive analysis of Human proteins And their disOrdered Segments - CHAOS). We implement internally consistent definitions of disordered regions, and annotate general characteristics such as cellular location, essentiality, post-translational modifications, and predicted pathogenicity. Further, we cross-reference to structure predictions from AlphaFold. We find that most human proteins contain at least one disordered region, predominantly located at the protein termini. IDRs are less hydrophobic, enriched in post-translational modifications, and mutations in IDRs are predicted to be less pathogenic than in non-IDRs. Additionally, we discovered that proteins residing in different cellular locations possess distinct disorder profiles. Finally, the predicted AlphaFold models of proteins in CHAOS suggest that disordered regions and proteins are often predicted to adopt secondary structure. Hereby we enhance the visibility and understanding of intrinsic disorder in human proteins. Key messagesO_LIFour out of five human proteins contain one or more intrinsically disordered regions (IDRs). C_LIO_LIHalf of the IDRs are located at protein termini, but three quarters of all human proteins contain a terminal IDR. C_LIO_LIThe amount and location of disordered regions differs throughout cellular compartments. C_LIO_LIOne in five missense mutations in IDRs are likely pathogenic. C_LIO_LIAlphaFold predicts secondary structure elements within intrinsically disordered regions and fully disordered proteins. C_LI
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- The structural coverage of the human proteome before and after AlphaFold 96%
- Protein prediction models support widespread post-transcriptional regulation of protein abundance by interacting partners 96%
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 95%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.