Which Variable Should Be Dependent in Phylogenetic Generalized Least Squares Regression Analysis
Chen, Z.-L.; Guo, H.-J.; Niu, D.-K.
Show abstract
Phylogenetic generalized least squares (PGLS) regression is widely used to examine evolutionary associations while accounting for phylogenetic non-independence, but it requires designating one trait as the dependent variable and the other as the independent variable. When causal relationships between traits are unclear, choosing which trait serves as the dependent variable becomes a practical concern. While studying the evolutionary relationship between bacterial growth rate and CRISPR-Cas content, we noticed that switching the roles of dependent and independent variables in PGLS analyses could lead to inconsistent conclusions in a substantial proportion of cases. To confirm this observation, we conducted 16,000 simulations of trait evolution along binary trees with 100 terminal nodes under different evolutionary models. PGLS regressions using Pagels {lambda} model were applied to each simulation, and the results consistently showed that swapping the dependent and independent variables can lead to inconsistent outcomes. We evaluated seven potential criteria for selecting the dependent variable, including log-likelihood, AIC, R2, p-value, Pagels {lambda}, Blombergs K, and the estimated {lambda} in the Pagels {lambda} model. Among these, Pagels {lambda}, Blombergs K, and the estimated {lambda} performed equally well and outperformed the others in selecting the dependent variable, providing a reliable basis for PGLS analyses when causal direction between traits is unclear.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A probabilistic model for indel evolution: differentiating insertions from deletions 95%
- ModelTeller: model selection for optimal phylogenetic reconstruction using machine learning 95%
- Benchmark of Differential Gene Expression Analysis Methods for Inter-species RNA-Seq Data using a Phylogenetic Simulation Framework 95%
Similar papers in this journal
- A variable-rate quantitative trait evolution model using penalized-likelihood 94%
- Evaluating probabilistic programming and fast variational Bayesian inference in phylogenetics 93%
- Distributions of extinction times from fossil ages and tree topologies: the example of mid-Permian synapsid extinctions 92%
Similar papers in this journal
- Inferring state-dependent diversification rates using approximate Bayesian computation (ABC) 96%
- Performance of a phylogenetic independent contrast method and an improved pairwise comparison under different scenarios of trait evolution after speciation and duplication 95%
- A semi-variance approach to visualising phylogenetic autocorrelation 95%
Similar papers in this journal
- PsiPartition: Improved Site Partitioning for Genomic Data by Parameterized Sorting Indices and Bayesian Optimization 93%
- Signatures of relaxed selection in the CYP8B1 gene of birds and mammals 91%
- Extant Sequence Reconstruction: The accuracy of ancestral sequence reconstructions evaluated by extant sequence cross-validation 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.