Identifying summary statistics for approximate Bayesian computation in a phylogenetic island biogeography model
Xie, S.; Valente, L.; Etienne, R.
Show abstract
Estimation of parameters of evolutionary island biogeography models, such as colonization and diversification rates, is important for a better understanding of island systems. A popular statistical inference framework is likelihood-based estimation of parameters using island species richness and phylogenetic data. Likelihood approaches require that the likelihood can be computed analytically or numerically, but with the increasing complexity of island biogeography models, this is often unfeasible. Simulation-based estimation methods may then be a promising alternative. One such method is approximate Bayesian computation (ABC), which compares summary statistics of the empirical data with the output of model simulations. However, ABC demands the definition of summary statistics that sufficiently describe the data, which is yet to be explored in island biogeography. Here, we propose a set of summary statistics and use it in an ABC framework for the estimation of parameters of an island biogeography model, DAISIE (Dynamic Assembly of Island biota through Speciation, Immigration and Extinction). For this model, likelihood-based inference is possible, which gives us the opportunity to assess the performance of the summary statistics. DAISIE currently only allows maximum likelihood estimation (MLE), so we additionally develop a likelihood-based Bayesian inference framework using Markov Chain Monte Carlo (MCMC) to enable comparison with the ABC results (i.e., making the same assumptions on prior distributions). We simulated phylogenies of island communities subject to colonization, speciation, and extinction using the DAISIE simulation model and compared the estimated parameters using the three inference approaches (MLE, MCMC and ABC). Our results show that the ABC algorithm performs well in estimating colonization and diversification rates, except when the species richness or amount of phylogenetic information from an island are low. We find that compared to island species diversity statistics, summary statistics that make use of phylogenetic and temporal patterns (e.g., the number of species through time) significantly improve ABC inference accuracy, especially in estimating colonization and anagenesis rates, as well as making inference converge considerably faster and perform better under the same number of iterations. Island biogeography is rapidly developing new simulation models that can explain the complexity of island biodiversity, and our study provides a set of informative summary statistics that can be used in island biogeography studies for which likelihood-based inference methods are not an option.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The robustness of a simple dynamic model of island biodiversity to geological and eustatic change 94%
- A shift in the host web occupancy of dew-drop spiders associated with genetic divergence in the Southwest Pacific 93%
- A joint distribution framework to improve presence-only species distribution models by exploiting opportunistic surveys 92%
Similar papers in this journal
- Using a GTR+Γ substitution model for dating sequence divergence when stationarity and time-reversibility assumptions are violated 93%
- Building alternative consensus trees and supertrees using k-means and Robinson and Foulds distance 93%
- Data-driven speciation tree prior for better species divergence times in calibration-poor molecular phylogenies 92%
Similar papers in this journal
- Distributions of extinction times from fossil ages and tree topologies: the example of mid-Permian synapsid extinctions 92%
- Evaluating probabilistic programming and fast variational Bayesian inference in phylogenetics 92%
- Geographic potential of the world largest hornet, Vespa mandarinia Smith (Hymenoptera: Vespidae), worldwide and particularly in North America 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.