Back

A Numerical Representation and Classification of Codons to Investigate Codon Alternation Patterns during Genetic Mutations on Disease Pathogenesis

Sengupta, A.; Pal Choudhury, P.; Chakraborty, S.; Roy, S.; Das, J. K.; Mallick, D.; Jana, S. S.

2020-03-03 bioinformatics
10.1101/2020.03.02.971036 bioRxiv
Show abstract

Alteration of amino acids is possible due to mutation in codons that could have potential reasons to occur disease. Single nucleotide substitutions (SNS) in genetic codon thus have prime importance for their ability to occur mutations that may be deleterious indeed. Effective mutation analysis can help to predict the fate of the diseased individual which can be validated later by in-vitro experiments. Hence in this present study, we try to investigate the codon alteration patterns and their impact during mutation for the genes known to be responsible for a particular disease. We use a numerical representation of four nucleotides based on the number of hydrogen bonds in their chemical structures and make a classification of 64 codons as well as corresponding 20 amino acids into three different classes (Strong, Weak and Transitional). The entire analysis has been carried out based on these classifications. For our current study, we consider two neurodegenerative diseases, Parkinsons disease, and Glaucoma. Several evidences claim similarities between both the diseases but proper pathogenetic factors are still unknown. The analysis reveals that the strong class of codons is highly mutated followed by the weak and transitional class. We observe that most of the mutations occur in the first or second positions in the codon rather than the third and mutations that occurred at the second place of codons are majorly deleterious. In most cases, the change in the determinative degree of codon due to mutation is directly proportional to the physical density property. Furthermore, we derive a determinative degree of five wild-type amino acid sequences, which can help biologists to understand the evolutionary relationship among them based on amino acid occurrence frequencies in proteins. In this regard we proposed an alignment-free method SSADDA (Sequence Similarity Analysis using Determinative Degree of Amino acid). Thus, our scheme gives a more microscopic and alternative representation of the existing codon table that helps in deciphering interesting codon alteration patterns during mutations in disease pathogenesis.

Matching journals

The top 10 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.