Leveraging a large language model to predict protein phase transition: a physical, multiscale and interpretable approach
Frank, M.; Ni, P.; Jensen, M.; Gerstein, M. B.
Show abstract
Protein phase transitions (PPTs) from the soluble state to a dense liquid phase (forming droplets via liquid-liquid phase separation) or to solid aggregates (such as amyloids) play key roles in pathological processes associated with age-related diseases such as Alzheimers disease. Several computational frameworks are capable of separately predicting the formation of droplets or amyloid aggregates based on protein sequences, yet none have tackled the prediction of both within a unified framework. Recently, large language models (LLMs) have exhibited great success in protein structure prediction; however, they have not yet been used for PPTs. Here, we fine-tune a LLM for predicting PPTs and demonstrate its usage in evaluating how sequence variants affect PPTs, an operation useful for protein design. In addition, we show its superior performance compared to suitable classical benchmarks. Due to the "black-box" nature of the LLM, we also employ a classical random forest model along with biophysical features to facilitate interpretation. Finally, focusing on Alzheimers disease-related proteins, we demonstrate that greater aggregation is associated with reduced gene expression in AD, suggesting a natural defense mechanism. Significance StatementProtein phase transition (PPT) is a physical mechanism associated with both physiological processes and age-related diseases. We present a modeling approach for predicting the protein propensity to undergo PPT, forming droplets or amyloids, directly from its sequence. We utilize a large language model (LLM) and demonstrate how variants within the protein sequence affect PPT. Because the LLM is naturally domain-agnostic, to enhance interpretability, we compare it with a classical knowledge-based model. Furthermore, our findings suggest the possible regulation of PPT by gene expression and transcription factors, hinting at potential targets for drug development. Our approach demonstrates the usefulness of fine-tuning a LLM for downstream tasks where only small datasets are available.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 94%
- CINS: Cell Interaction Network inference from Single cell expression data 94%
- Protein structural features predict responsiveness to pharmacological chaperone treatment for three lysosomal storage disorders 93%
Similar papers in this journal
- In silico characterization of mechanisms positioning costimulatory and checkpoint complexes in immune synapses 92%
- Evolutionary analysis reveals the role of a non-catalytic domain of peptidyl arginine deiminase 2 in transcriptional regulation 92%
- The Division of Amyloid Fibrils - Systematic comparison of fibril fragmentation stability by linking theory with experiments 91%
Similar papers in this journal
- Intrinsically disordered regions that drive phase separation form a robustly distinct protein class 96%
- Beta Turn Propensity and a Model Polymer Scaling Exponent Identify Disordered Proteins that Phase Separate 96%
- The bacterial chaperone CsgC inhibits functional amyloid CsgA formation by promoting the intrinsically disordered pre-nuclear state 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.