Back

A Systematic Evaluation of Protein Phase Separation Predictors Across Diverse Protein Landscapes

Gilroy, K. E.; Barr, J. N.

2026-01-22 bioinformatics
10.64898/2026.01.20.700394 bioRxiv
Show abstract

BackgroundLiquid-liquid phase separation (LLPS) plays a central role in cellular regulation, with its dysregulation linked to numerous disease outcomes. LLPS has also being increasingly implicated as critical in various biological contexts, with a key example being virus replication. These findings have driven the development of numerous computational predictors to screen and identify phase-separating proteins from sequence and/or structural models. Despite the growing need for these tools, their comparative performance across diverse biological contexts remains incompletely understood, complicating tool selection and result interpretation. ResultsA systematic comparative analysis of nine LLPS prediction algorithms was conducted using multiple curated datasets comprising both LLPS-positive (LLPS+) and -negative (LLPS-) proteins. The datasets span a range of biologically relevant scenarios, including intrinsically disordered proteins, folded proteins, proteins with LLPS-abolishing variants, benchmark datasets, and viral proteins. Substantial variability in predictive performance was observed across algorithms when assessing proteins of different structural classes. Numerous showed reduced accuracy in distinguishing LLPS+ and LLPS- folded proteins. Sensitivity to the impact of small LLPS-abolishing mutations also varied markedly between predictors, with structure-informed algorithms generally outperforming most sequence-based predictors. Similar decreases in predictive capability were also observed across several algorithms when analysing viral proteins, which are often-under-represented in predictor training datasets. ConclusionsThese results demonstrate that LLPS predictor performance is strongly context dependent, leading to different predictors being optimal for different biological questions. For overall protein assessment, DeePhase and Molphase provided the most consistently accurate predictions, being the least impacted by structural bias. For assessing the impact of small mutations on LLPS propensity, PSPHunter, a built-for-purpose algorithm, reliably predicts mutation impacts with structure-informed algorithms PSPire and PICNIC also providing strong insight. Across all evaluated datasets, the findings highlight the need for well-benchmarked training and testing data that encompasses a broad and diverse range of protein classes.

Published in Computational and Structural Biotechnology Journal (predicted rank #4) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.