Back

Conformational ensembles of the human intrinsically disordered proteome: Bridging chain compaction with function and sequence conservation

Tesei, G.; Trolle, A. I.; Jonsson, N.; Betz, J.; Pesce, F.; Johansson, K. E.; Lindorff-Larsen, K.

2023-05-08 biophysics
10.1101/2023.05.08.539815 bioRxiv
Show abstract

Intrinsically disordered proteins and regions (collectively IDRs) are pervasive across proteomes in all kingdoms of life, help shape biological functions, and are involved in numerous diseases. IDRs populate a diverse set of transiently formed structures, yet defy commonly held sequence-structure-function relationships. Recent developments in protein structure prediction have led to the ability to predict the three-dimensional structures of folded proteins at the proteome scale, and have enabled large-scale studies of structure-function relationships. In contrast, knowledge of the conformational properties of IDRs is scarce, in part because the sequences of disordered proteins are poorly conserved and because only few have been characterized experimentally. We have developed an efficient model to generate conformational ensembles of IDRs, and thereby to predict their conformational properties from sequence only. Here, we applied this model to simulate all IDRs of the human proteome. Examining conformational ensembles of 29,998 IDRs, we show how chain compaction is correlated with cellular function and localization, including in different types of biomolecular condensates. We train a model to predict compaction from sequence and use this to show conservation of structural properties across orthologs. Our results recapitulate observations from previous studies of individual protein systems, and enable us to study the relationship between sequence, conservation, conformational ensembles, biological function and disease variants at the proteome scale.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.