Supervised learning of protein variant effects across large-scale mutagenesis datasets
Schulze, T. K.; Blaabjerg, L. M.; Cagiada, M.; Lindorff-Larsen, K.
Show abstract
The increasing availability of data from multiplexed assays of variant effects (MAVEs) enables supervised model training against large quantities of experimental data to learn sequence-function relationships. Variant effect scores from MAVEs can, however, be influenced by the experimental method to create experiment-to-experiment differences in the mapping from molecular level-variant effects to MAVE readout, which presents a challenge for supervised learning across datasets. We here propose a framework for performing supervised learning with MAVE data that takes the influence of the experimental protocol into account, thus enabling variant effects to be learned across datasets produced in independent experiments. We apply the framework to train a model against variant effect scores collected with VAMP-seq, a MAVE technique that quantifies the steady-state cellular abundance of protein variants. We show that mapping variant abundance to VAMP-seq readout in a dataset-specific manner during model training improves the learned abundance model and moreover allows the learned model to predict variant effects on an interpretable scale. Our work highlights the importance of validating MAVE results with low-throughput methods to facilitate MAVE score interpretation and supervised model training.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Overcoming the design, build, test (DBT) bottleneck for synthesis of nonrepetitive protein-RNA binding cassettes for RNA applications 94%
- Cell-type-specific co-expression inference from single cell RNA-sequencing data 94%
- Understanding epistatic networks in the B1 -lactamases through coevolutionary statistical modeling and deep mutational scanning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.