Improved Phylogenetic Posterior Estimation through Regularised Conditional Clade Distributions
Yang, Z.; Klawitter, J.; Bouckaert, R.; Drummond, A.
Show abstract
Bayesian phylogenetic inference uses Markov chain Monte Carlo sampling to estimate the posterior distribution of phylogenetic trees. However, the complex geometry of treespace makes these distributions difficult to characterise. Traditional approaches often summarise posterior samples into a single point estimate, discarding much of the information contained in the full distribution. Conditional clade distributions (CCDs) address this by providing a tractable model of the full tree distribution, but existing parameterisations exhibit complementary strengths. We introduce a new family of models called regularised conditional clade distributions (regCCD), which apply regularisation to balance sample fidelity against overfitting, thereby combining the strengths of existing parameterisations. We show that regCCD outperforms existing models in capturing the posterior distribution and provides better point estimates than its underlying model. Furthermore, we provide an efficient procedure for selecting the optimal regularisation parameter for a given posterior tree set. Finally, we demonstrate the application of regCCD in assessing whether different phylogenetic models or data produce statistically distinguishable tree distributions.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Adaptive Tree Proposals for Bayesian Phylogenetic Inference 96%
- Using Parsimony-Guided Tree Proposals to Accelerate Convergence in Bayesian Phylogenetic Inference 96%
- ConvexML: Fast and accurate branch length estimation under irreversible mutation models, illustrated through applications to CRISPR/Cas9-based lineage tracing 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.