A Transformer Based Machine Learning of Molecular Grammar Inherent in Proteins Prone to Liquid Liquid Phase Separation
Wasim, A.; Mondal, J.
Show abstract
Understanding the molecular grammar that governs protein phase separation is essential for advancements in bioinformatics and protein engineering. This study leverages Generative Pre-trained Transformer (GPT)-based Protein Language Models (PLMs) to decode the complex grammar of proteins prone to liquid-liquid phase separation (LLPS). We trained three distinct GPT models on datasets comprising amino acid sequences with varying LLPS propensities: highly predisposed (LLPS+ GPT), moderate (LLPS-GPT), and resistant (PDB* GPT). As training progressed, the LLPS-prone model began to learn embeddings that were distinct from those in LLPS-resistant sequences. These models generated 18,000 protein sequences ranging from 20 to 200 amino acids, which exhibited low similarity to known sequences in the SwissProt database. Statistical analysis revealed subtle but significant differences in amino acid occurrence probabilities between sequences from LLPS-prone and LLPS-resistant models, suggesting distinct molecular grammar underlying their phase separation abilities. Notably, sequences from LLPS+ GPT showed fewer aromatic residues and a higher fraction of charge decoration. Short peptides (20-25 amino acids) generated from LLPS+ GPT underwent computational and wet-lab validation, demonstrating their ability to form phase-separated states in vitro. The generated sequences enriched the existing database and enabled the development of a robust classifier that accurately distinguishes LLPS-prone from non-LLPS sequences. This research marks a significant advancement in using computational models to explore and engineer the vast protein sequence space associated with LLPS-prone proteins.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Accurate Conformation Sampling via Protein Structural Diffusion 97%
- From Signal to Symphony: Exploring 2D Sequence Representations for Protein Function Prediction 96%
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 96%
Similar papers in this journal
- AlphaMut: a deep reinforcement learning model to suggest helix-disrupting mutations 97%
- Thermal Adaptation of Cytosolic Malate Dehydrogenase Revealed by Deep Learning and Coevolutionary Analysis 96%
- Accurate and Rapid Prediction of Protein pKa: Protein Language Models Reveal the Sequence-pKa Relationship 95%
Similar papers in this journal
- Rationalize the Functional Roles of Protein-Protein Interactions in Targeted Protein Degradation by Kinetic Monte-Carlo Simulations 95%
- Effect of Mutations on Smlt1473 Binding to Various Substrates Using Molecular Dynamics Simulations 95%
- Unraveling the Molecular Complexity of N-TerminusHuntingtin Oligomers: Insights into Polymorphic Structures 95%
Similar papers in this journal
- A Topological Data Analytic Approach for Discovering Biophysical Signatures in Protein Dynamics 96%
- Transferable deep generative modeling of intrinsically disordered protein conformations 96%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.