Removing bias in sequence models of protein fitness
Shaw, A. Y.; Spinner, H. B.; Gurev, S.; Shin, J.-E.; Rollins, N.; Marks, D. S.
Show abstract
Unsupervised sequence models for protein fitness have emerged as powerful tools for protein design in order to engineer therapeutics and industrial enzymes, yet they are strongly biased towards potential designs that are close to their training data. This hinders their ability to generate functional sequences that are far away from natural sequences, as is often desired to design new functions. To address this problem, we introduce a de-biasing approach that enables the comparison of protein sequences across mutational depths to overcome the extant sequence similarity bias in natural sequence models. We demonstrate our methods effectiveness at improving the relative natural sequence model predictions of experimentally measured variant functions across mutational depths. Using case studies proteins with very low functional percentages further away from the wild type, we demonstrate that our method improves the recovery of top-performing variants in these sparsely functional regimes. Our method is generally applicable to any unsupervised fitness prediction model, and for any function for any protein, and can thus easily be incorporated into any computational protein design pipeline. These studies have the potential to develop more efficient and cost-effective computational methods for designing diverse functional proteins and to inform underlying experimental library design to best take advantage of machine learning capabilities.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- RP3Net: a deep learning model for predicting recombinant protein production in Escherichia coli 95%
- CaLMPhosKAN: Prediction of General Phosphorylation Sites in Proteins via Fusion of Codon Aware Embeddings with Amino Acid Aware Embeddings and Wavelet-based Kolmogorov Arnold Network 95%
- Coevolution-based prediction of protein-protein interactions in polyketide biosynthetic assembly lines 95%
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 96%
- Data-efficient protein mutational effect prediction with weak supervision by molecular simulation and protein language models 94%
- Cracking the black box of deep sequence-based protein-protein interaction prediction 94%
Similar papers in this journal
- Identifying promising sequences for protein engineering using a deep Transformer Protein Language Model 96%
- Improved prediction of stabilizing mutations in proteins by incorporation of mutational effects on ligand binding 95%
- Identification of biochemically neutral positions in liver pyruvate kinase 95%
Similar papers in this journal
- ProteinGLUE: A multi-task benchmark suite for self-supervised protein modeling. 95%
- PIPENN-EMB: ensemble net and protein embeddings generalise protein interface prediction beyond homology 95%
- Evaluating the Significance of Embedding-Based Protein Sequence Alignment with Clustering and Double Dynamic Programming for Remote Homology 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.