Inferring Protein Variant Impacts Across Contexts
Rasoulzadeh Hosseini, A.; Senguttuvan, V.; van Loggerenberg, W.; Border, R.; Roth, F. P.
Show abstract
Multiplexed assays of variant effects (MAVEs) measure the functional impact of many protein sequence variants in parallel, potentially covering all possible single amino acid substitutions. Unlike current computational variant effect predictors, MAVEs can reveal the effects of variants under different genetic and environmental contexts. However, whereas the space of possible contexts is effectively infinite, contextual MAVE studies are limited by finite experimental budgets. To maximize coverage across contexts, one strategy is to carry out sub-saturation contextual MAVEs and then fill in the gaps via imputation. Here, we categorize and compare different imputation challenges, explore a collection of multi-context imputation solutions, including linear mixed-effects models, random forests, and autoencoders, and provide insight into how best to proceed for a given imputation task. We find that the optimal method depends on the imputation task and how densely the contexts have been measured. More flexible models excel when measurements are plentiful, whereas the simplest models prove most reliable when measurements are sparse. However, the simple source-to-target regression models, although well suited to imputing scores for variants measured in the source context, cannot impute scores for variants that were not measured in either context. This is a major limitation when both maps are sparsely measured. We provide a conceptual framework and an initial evaluation of multi-context imputation methods that can extend the scope of large-scale studies of context-dependent variant effects.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The simplicity of protein sequence-function relationships 94%
- MutPred2: inferring the molecular and phenotypic impact of amino acid variants 94%
- Single-Cell Omics for Transcriptome CHaracterization (SCOTCH): isoform-level characterization of gene expression through long-read single-cell RNA sequencing 94%
Similar papers in this journal
- An open-source platform to distribute and interpret data from multiplexed assays of variant effect 94%
- Fine-tuning sequence-to-expression models onpersonal genome and transcriptome data 94%
- EvoAug: improving generalization and interpretability of genomic deep neural networks with evolution-inspired data augmentations 94%
Similar papers in this journal
Similar papers in this journal
- A tissue-aware machine learning framework enhances the mechanistic understanding and genetic diagnosis of Mendelian and rare diseases 93%
- BaCoN (Balanced Correlation Network) improves prediction of gene buffering 92%
- Interpretable deep generative ensemble learning for single-cell omics with Hydra 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.