Back

Identifying key amino acid types that distinguish paralogous proteins using Shapley value based feature subset selection

Machingal, P.; Busi, R.; Hemachandra, N.; Balaji, P. V.

2024-11-01 bioinformatics
10.1101/2024.04.26.591291 bioRxiv
Show abstract

We view a protein as the composite of the standard 20 amino acids, ignoring their order in the protein sequence. With this view, we try to identify the important amino acid types that distinguish pairs of paralogous proteins, thereby playing a role in their functional difference. Using only the amino acid composition (AAC) as features and a linear classifier, we find that many pairs of paralogous protein families can be classified accurately. Next, we use an existing Shapley value-based feature subset selection algorithm, SVEA, to identify the important amino acid types that distinguish a pair of paralogous proteins. The SVEA algorithm assigns a score, Shapley value, to each feature, amino acid type, based on its contribution to the classifiers training error. We identify the important distinguishing amino acid types as those whose Shapley value exceeds a data-driven threshold. We refer to these as the amino acid feature subset (AFS). We find that many paralog pairs can still be accurately classified using only the AFS composition. We partition AFS based on the classifier weights to infer class-wise amino acid importance. We verify whether the identified AFS amino acids indeed play a role in the functional difference of the paralog pairs using various methods - multiple sequence alignment, 3D structure analysis, and supporting evidence from biology literature. We also discuss some consistencies observed in the Shapley value based ranking and the AFS when comparing the AFS of two different but related paralog pairs. We demonstrate the results for 15 pairs of paralogous proteins.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.