VEFill: a model for accurate and generalizable deep mutational scanning score imputation across protein domains
Polunina, P. V.; Maier, W.; Rubin, A. F.
Show abstract
BackgroundDeep Mutational Scanning (DMS) assays can systematically assess the effects of amino acid substitutions on protein function. While DMS datasets have been generated for many targets, they often suffer from incomplete variant coverage due to technical constraints, limiting their utility in variant interpretation and downstream analyses. ResultsWe developed VEFill, a gradient boosting model for imputing missing DMS scores across protein domains. VEFill is trained on the Human Domainome 1 dataset, a large, standardized set of DMS experiments using a uniform stability-based assay, and integrates a broad set of additional biologically informative features including ESM-1v sequence embeddings, evolutionary conservation (EVE scores), amino acid substitution matrices, and physicochemical descriptors. The model achieved robust predictive performance (R2 = 0.64, Pearson r = 0.80). It also demonstrated reliable generalization to unseen proteins in other stability-based datasets, while showing weaker performance on activity-based assays. Per-protein models further confirmed VEFills effectiveness under limited-data conditions. A reduced two-feature version using only ESM-1v embeddings and mean DMS scores performed comparably to the full model, suggesting a computationally efficient alternative. However, true zeroshot prediction without positional context remains a challenge, particularly for functionally complex proteins. ConclusionsVEFill offers an interpretable, scalable framework for DMS score imputation, especially effective in stability-focused and sparse-data settings. It enables systematic mutation prioritization and may support the design of efficient experimental libraries for variant effect studies.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Beyond the Leaderboard: Leveraging Predictive Modeling for Protein-Ligand Insights and Discovery 96%
- Expert-guided protein Language Models enable accurate and blazingly fast fitness prediction 96%
- Unsupervised protein embeddings outperform hand-crafted sequence and structure features at predicting molecular function 96%
Similar papers in this journal
Similar papers in this journal
- Controllable Protein Design via Autoregressive Direct Coupling Analysis Conditioned on Principal Components 95%
- Paraplume: A fast and accurate paratope prediction method provides insights into repertoire-scale binding dynamics 94%
- Nucleotide context models outperform protein language models for predicting antibody affinity maturation 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.