Identifying key amino acid types that distinguish paralogous proteins using Shapley value based feature subset selection
Machingal, P.; Busi, R.; Hemachandra, N.; Balaji, P. V.
Show abstract
We view a protein as the composite of the standard 20 amino acids, ignoring their order in the protein sequence. With this view, we try to identify the important amino acid types that distinguish pairs of paralogous proteins, thereby playing a role in their functional difference. Using only the amino acid composition (AAC) as features and a linear classifier, we find that many pairs of paralogous protein families can be classified accurately. Next, we use an existing Shapley value-based feature subset selection algorithm, SVEA, to identify the important amino acid types that distinguish a pair of paralogous proteins. The SVEA algorithm assigns a score, Shapley value, to each feature, amino acid type, based on its contribution to the classifiers training error. We identify the important distinguishing amino acid types as those whose Shapley value exceeds a data-driven threshold. We refer to these as the amino acid feature subset (AFS). We find that many paralog pairs can still be accurately classified using only the AFS composition. We partition AFS based on the classifier weights to infer class-wise amino acid importance. We verify whether the identified AFS amino acids indeed play a role in the functional difference of the paralog pairs using various methods - multiple sequence alignment, 3D structure analysis, and supporting evidence from biology literature. We also discuss some consistencies observed in the Shapley value based ranking and the AFS when comparing the AFS of two different but related paralog pairs. We demonstrate the results for 15 pairs of paralogous proteins.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 97%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 96%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 95%
Similar papers in this journal
- DELPHI: accurate deep ensemble model for protein interaction sites prediction 96%
- Patch-DCA: Improved Protein Interface Prediction by utilizing Structural Information and Clustering DCA scores 95%
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 95%
Similar papers in this journal
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 96%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 96%
- Alignment of virus-host protein-protein interaction networks by integer linear programming: SARS-CoV-2 96%
Similar papers in this journal
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 96%
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 96%
- A method for predicting linear and conformational B-cell epitopes in an antigen from its primary sequence 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.