Data-efficient protein mutational effect prediction with weak supervision by molecular simulation and protein language models
Deguchi, T.; Kurumida, Y.; Iida, S.; Kobayashi, K.; Saito, Y.
Show abstract
Machine learning-based protein mutational effect prediction is widely used in protein engineering and pathogenicity prediction, but training data scarcity remains a major challenge due to high costs of experimental measurements. A previous study proposed data augmentation using computational estimates by molecular simulation. However, this approach has been limited to predicting mutational effects on thermostability. Here, we present a new data augmentation method that combines molecular simulation with zero-shot prediction computed by protein language models. These computational estimates serve as "weak" training data to supplement experimental training data. Our method dynamically adjusts the weight and inclusion of weak training data based on available experimental training data. This reduces potential negative impacts of weak training data while extending applicability to diverse protein properties such as binding affinity and enzymatic activity. Benchmark tests demonstrate that our method improves prediction accuracy particularly when experimental training data are scarce. These results indicate the capability of our approach to advance protein engineering and pathogenicity prediction in small data regimes.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- VoroCNN: Deep convolutional neural network built on 3D Voronoi tessellation of protein structures 96%
- Enhancing predictions of protein stability changes induced by single mutations using MSA-based Language Models 95%
- CONSTRUCT: an algorithmic tool for identifying functional or structurally important regions in protein tertiary structure 95%
Similar papers in this journal
Similar papers in this journal
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 95%
- Accurate Conformation Sampling via Protein Structural Diffusion 95%
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.