Variant characterization in the intrinsically disordered human proteome
Hubrich, D.; Alvarado Valverde, J.; Lee, C. Y.; Djokic, M.; Welzel, M.; Hintz, K.; Luck, K.
Show abstract
Variant effect prediction remains a key challenge to resolve in precision medicine. Sophisticated computational models that exploit sequence conservation and structure are increasingly successful in the characterization of missense variants in folded protein regions. However, 37% of all annotated missense variants reside in 25% of the proteome that is intrinsically disordered, lacking positional sequence conservation and stable structures. To significantly advance variant effect prediction in disordered protein regions, we combined sequence pattern searches with AlphaFold and experiments to structurally annotate 1,300 protein-protein interactions with interfaces mediated by short disordered motifs binding to folded domains in partner proteins. These interfaces were selected based on their overlap with uncertain missense variants enabling reliable prediction of deleterious effects of 1,187 uncertain variants in disordered protein regions. Extensive experimental efforts validate predicted interfaces and deleterious variant effects that were predicted as benign by AlphaMissense. This study demonstrates how critical structural information on protein interaction interfaces is for variant effect prediction especially in disordered protein regions and provides a clear avenue towards its system-wide implementation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Assembly defects of the human tRNA splicing endonuclease contribute to impaired pre-tRNA processing in pontocerebellar hypoplasia 98%
- Architecture and regulation of filamentous humancystathionine beta-synthase 98%
- A cell type-aware framework for nominating non-coding variants in Mendelian regulatory disorders 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.