Fitness translocation: improving variant effect prediction with biologically-grounded data augmentation
Mialland, A.; Fukunaga, S.; Katsuki, R.; Dong, Y.; Yamaguchi, H.; Saito, Y.
Show abstract
Predicting the functional effects of protein variants (variant effect prediction) is essential in protein engineering but remains challenging due to the scarcity of fitness data for training prediction models. To address this limitation, we introduce a data augmentation strategy called fitness translocation, which leverages variant fitness data from homologous proteins to enhance prediction models for a target protein. Using embeddings from protein language models, our method computes the differences between the homologs wild type and its variants, which are applied to the target wild type to generate its synthetic variants in the embedding space. We evaluate this approach on three protein families: IGPS, GFP, and SARS-CoV-2 spike proteins, under various prediction models and training data sizes. Fitness translocation consistently improves prediction accuracy, especially under limited training data. Moreover, accuracy improvement is observed even between remote homologs with sequence identity as low as 35%. These results highlight the potential of data-efficient protein engineering by reusing fitness data previously accumulated in homologs. The code is available at https://github.com/adrienmialland/ProtFitTrans.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
- Leveraging Permutation Testing to Assess Confidence in Positive-Unlabeled Learning Applied to High-Dimensional Biological Datasets 95%
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 95%
- Challenges in predicting protein-protein interactions of understudied viruses: Arenavirus-Human interactions 95%
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 94%
Similar papers in this journal
- In-Pero: Exploiting deep learning embeddings of protein sequences to predict the localisation of peroxisomal proteins 95%
- Protein-protein interaction prediction for targeted protein degradation 95%
- DSResSol: A sequence-based solubility predictor created with Dilated Squeeze Excitation Residual Networks 94%
Similar papers in this journal
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 96%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 95%
- Development of an absolute assignment predictor for triple-negative breast cancer subtyping using machine learning approaches 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.