Interpretable Prediction of Phase Separation and Disease Variant Effects in Intrinsically Disordered Regions
Zhao, M.; Kumar, S.
Show abstract
Coding mutations within intrinsically disordered regions (IDRs) of proteins are increasingly implicated in human diseases yet remain poorly interpreted by conventional variant-effect predictors that rely on structural stability and conservation-based metrics. Quantifying disruption of IDR-mediated liquid-liquid phase separation (LLPS) offers a biophysically principled approach to interpreting the pathogenic impact of such variants. However, existing LLPS predictors suffer from training biases toward self-separating proteins, show limited performance on partner- dependent phase separation, and often lack interpretability for variant prioritization. We present an interpretable ensemble machine-learning framework that integrates protein language model embeddings of sequence and predicted structure to predict LLPS propensity and classify proteins as self-separating or partner-dependent. Our two-step classifiers outperform existing methods on independent benchmark datasets, with the largest gains for partner-dependent LLPS proteins. Beyond classification, our framework identifies critical phase-separating regions and quantifies mutation-induced perturbations in LLPS. Applied to disease-associated variant databases, we found that pathogenic mutations are enriched in predicted phase-separating regions and frequently perturb LLPS propensity scores, implicating mutation-induced LLPS dysregulation as a potential pathogenic mechanism for numerous diseases. Overall, our framework provides an accurate, interpretable approach for identifying phase-separating proteins and linking aberrant phase- separation behavior to disease pathogenesis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Decoding Missense Variants by Incorporating Phase Separation via Machine Learning 97%
- Generalizable and scalable protein stability prediction with rewired protein generative models 97%
- PreMode predicts mode-of-action of missense variants by deep graph representation learning of protein sequence and structural context 96%
Similar papers in this journal
- hu.MAP3.0: Atlas of human protein complexes by integration of > 25,000 proteomic experiments 96%
- A tissue-aware machine learning framework enhances the mechanistic understanding and genetic diagnosis of Mendelian and rare diseases 96%
- PIFiA: Self-supervised Approach for Protein Functional Annotation from Single-Cell Imaging Data 96%
Similar papers in this journal
- Integrative, high-resolution analysis of single cell gene expression across experimental conditions with PARAFAC2-RISE 95%
- Dysregulation of the secretory pathway connects Alzheimer's disease genetics to aggregate formation 95%
- Engineering of highly active and diverse nuclease enzymes by combining machine learning and ultra-high-throughput screening 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.