G-domain prediction across the diversity of G protein families
Sanghavi, H. M.; Rashmi, R.; Dasgupta, A.; Majumdar, S.
Show abstract
Guanine nucleotide binding proteins are characterized by a structurally and mechanistically conserved GTP-binding domain, indispensable for binding GTP. The G domain comprises of five adjacent consensus motifs called G boxes, which are separated by amino acid spacers of different lengths. Several G proteins, discovered over time, are characterized by diverse function and sequence. This sequence diversity is also observed in the G box motifs (specifically the G5 box) as well as the inter-G box spacer length. The Spacers and Mismatch Algorithm (SMA) introduced in this study, can predict G-domains in a given G protein sequence, based on user-specified constraints for approximate G-box patterns and inter-box gaps in each G protein family. The SMA parameters can be customized as more G proteins are discovered and characterized structurally. Family-specific G box motifs including the less characterized G5 motif as well as G domain boundaries were predicted with higher precision. Overall, our analysis suggests the possible classification of G protein families based on family-specific G box sequences and lengths of inter-G box spacers. Significance StatementIt is difficult to define the boundaries of a G domain as well as predict G boxes and important GTP-binding residues of a G protein, if structural information is not available. Sequence alignment and phylogenetic methods are often unsuccessful, given the sequence diversity across G protein families. SMA is a unique method which uses approximate pattern matching as well as inter-motif separation constraints to predict the locations of G-boxes. It is able to predict all G boxes including the less characterized G5 motif which marks the carboxy-terminal boundary of a G domain. Thus, SMA can be used to predict G domain boundaries within a large multi-domain protein as long as the user-specified constraints are satisfied. ClassificationBiological Sciences/Biophysics and Computational Biology
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Conserved intramolecular networks in GDAP1 are closely connected to CMT-linked mutations and protein stability 94%
- Role of two modules controlling the interaction between SKAP1 and SRC kinasesComparison with SKAP2 architecture and consequences for evolution 93%
- Phylogenetic analysis of the MCL1 BH3 binding groove and rBH3 sequence motifs in the p53 and INK4 protein families 93%
Similar papers in this journal
- Bioinformatic Analysis of Structure and Function of LIM Domains of Human Zyxin Family Proteins 95%
- Evolutionary formation of a human de novo open reading frame from a non-primate non-coding genomic region that proceeds via biased random mutations 92%
- In silico investigation of the new UK (B.1.1.7) and South African (501Y.V2) SARS-CoV-2 variants with a focus at the ACE2-Spike RBD interface 92%
Similar papers in this journal
- Predicting human and viral protein variants affecting COVID-19 susceptibility and repurposing therapeutics 93%
- In Silico Analysis Predicting Effects of Deleterious SNPs of Human RASSF5 Gene on its Structure and Functions 92%
- LIM domain-wide comprehensive mutagenesis reveals the role of leucine in CSRP3 protein stability 91%
Similar papers in this journal
- A Mathematical Genomics Perspective on the Moonlighting Role of Glyceraldehyde-3-Phosphate Dehydrogenase (GAPDH) 94%
- Unveiling the Genetic Tapestry: Rare Disease Genomics of Spinal Muscular Atrophy and Phenylketonuria Proteins 94%
- The Distal-Proximal Relationships Among the Human Moonlighting Proteins: Evolutionary hotspots and Darwinian checkpoints 93%
Similar papers in this journal
- Definition and discovery of tandem SH3-binding motifs interacting with members of the p47phox-related protein family 92%
- CoRNeA: A pipeline to decrypt the protein-protein interaction from amino acid sequence information 91%
- Can we assume the gene expression profile as a proxy for signaling network activity? 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.