Back

Shedding light on the underlying characteristics of genomes using Kronecker model families of codon evolution

Zaheri, M.; Salamin, N.

2020-08-13 bioinformatics
10.1101/2020.08.12.247890 bioRxiv
Show abstract

The mechanistic models of codon evolution rely on some simplistic assumptions in order to reduce the computational complexity of estimating the high number of parameters of the models. This paper is an attempt to investigate how much these simplistic assumptions are misleading when they violate the nature of the biological dataset in hand. We particularly focus on three simplistic assumptions made by most of the current mechanistic codon models including: 1) only single substitutions between nucleotides within codons in the codon transition rate matrix are allowed. 2) mutation is homogenous across nucleotides within a codon. 3) assuming HKY nucleotide model is good enough at the nucleotide level. For this purpose, we developed a framework of mechanistic codon models, each model in the framework hold or relax some of the mentioned simplifying assumptions. Holding or relaxing the three simplistic assumptions results in total to eight different mechanistic models in the framework. Through several experiments on biological datasets and simulations we show that the three simplistic assumptions are unrealistic for most of the biological datasets and relaxing these assumptions lead to accurate estimation of evolutionary parameters such as selection pressure.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.