Back

Effects of residue substitutions on the cellular abundance of proteins

Schulze, T. K.; Lindorff-Larsen, K.

2024-09-24 biophysics Community evaluation
10.1101/2024.09.23.614650 bioRxiv
Show abstract

Multiplexed assays of variant effects (MAVEs) make it possible to measure the functional impact of all possible single amino acid residue substitutions in a protein in a single experiment. Combination of variant effect data from several such experiments provides the opportunity to conduct large-scale analyses of variant effect scores measured across proteins, but can be complicated by variations in the phenotypes that are probed across experiments. Thus, using variant effect datasets obtained with similar MAVE techniques can help reveal general rules governing the effects of amino acid variation for a single molecular phenotype. In this work, we accordingly combined data from six individual variant abundance by massively parallel sequencing (VAMP-seq) experiments and analysed a total of 31,614 variant effect scores reporting solely on the impact of single amino acid residue substitutions on the cellular abundance of proteins. Using our combined variant effect dataset, we derived and analysed a collection of amino acid substitution matrices describing the average impact on cellular abundance of all residue substitution types in different structural environments. We found that the substitution matrices predict the cellular abundance of protein variants with surprisingly high accuracy when given structural information only in the form of whether a residue is buried or exposed. We thus propose our substitution matrix-based predictions as strong baselines for future abundance model development.

Published in eLife (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.