Likelihood-based evaluation of character recoding schemes for phylogenetic analysis
Seo, T.-K.; Thorne, J. L.
Show abstract
Character recoding is a common practice in evolutionary studies. For example, phylogenies can be inferred from protein-coding DNA sequences via 61-state codon substitution models. However, often the inference is instead done by adopting 20-state amino acid replacement models that do not explicitly consider synonymous substitution. When there is substantial heterogeneity of amino acid frequencies among sites and/or among lineages, another sort of character recoding is sometimes performed. In these cases, one option is to reduce the state space of the model by placing each of the 20 amino acids into one of a relatively small number (e.g., 6) of groups of amino acids and then modeling only how group membership changes at a site over evolutionary time. Unfortunately, these kinds of character recoding schemes are prone to reducing the amount of available evolutionary information. Here, we provide a likelihood framework to statistically assess recoding schemes. Although we concentrate on the recoding of 61-state codon substitution models into 20-state amino acid replacement models, the general approach is also relevant to other recoding schemes such as those that recode 20-state models into 6-state models.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deformity Index: A semi-reference quality metric of phylogenetic trees based on their clades 92%
- Moderate amounts of epistasis are not evolutionarily stable in small populations 92%
- PsiPartition: Improved Site Partitioning for Genomic Data by Parameterized Sorting Indices and Bayesian Optimization 91%
Similar papers in this journal
Similar papers in this journal
- Gene Transfer based Phylogenetics: Analytical Expressions and Additivity via Birth/Death Theory 94%
- Consistency of SVDQuartets and Maximum Likelihood for Coalescent-based Species Tree Estimation 94%
- Biased estimates of phylogenetic branch lengths resulting from the discretised Gamma model of site rate heterogeneity 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.