Back

A machine learning framework for predicting and modulating condition-dependent protein phase separation

Bae, J.; Kang, M.; Lee, D.; Yoon, K.-J.; Jung, Y.

2025-12-29 bioinformatics
10.64898/2025.12.28.696755 bioRxiv
Show abstract

Protein phase separation is a fundamental process in organizing membraneless organelles and is implicated in a wide range of pathological conditions. Importantly, rather than being a static feature of specific proteins, phase separation is a condition-dependent phenomenon governed by environmental parameters, including protein concentration, temperature, and solvent composition. However, most existing machine learning models infer phase-separation propensity solely from amino-acid sequences, failing to capture these context-dependent behaviors. Here, we present LLPSense, a machine learning framework that integrates pre-trained protein language model embeddings with environmental parameters to achieve accurate, condition-aware predictions of protein phase separation. We demonstrate LLPSenses predictive power and utility through three key experimental demonstrations. First, the model revealed that SGTA, previously unrecognized as a phase-separating protein, exhibits complex, temperature-dependent reentrant phase behavior. Second, LLPSense accurately predicted mutations in -synuclein that either enhance or suppress phase separation, enabling systematic mapping of residues potentially relevant to Parkinsons disease. Third, using model-guided mutagenesis, we inverted the phase behavior of UBQLN4, shifting it from high-temperature to low-temperature separation. Collectively, LLPSense provides a robust computational tool for interrogating the condition-dependent landscape of protein phase separation, enabling mechanistic studies of disease-associated phase separation and the rational design of programmable condensates.

Published in Nature Communications (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.