Optimized Phylogenetic Eigenvector Regression Provides Robust and Flexible Trait Correlation Analysis under Diverse Evolutionary Conditions
Chen, Z.-L.; Niu, D.-K.
Show abstract
Phylogenetic eigenvector regression (PVR) is widely used in ecology and evolution by representing phylogenetic structure through separable eigenvectors. Despite this flexibility, its implementation faces three key challenges: (1) the selection of eigenvectors, (2) the reduced robustness of ordinary least-squares (OLS) regression under shift-like evolutionary heterogeneity, and (3) the applicability of conventional model complexity rules such as the "samples-per-variable (SPV) [≥] 10" guideline. Here, we propose an optimized PVR framework that addresses these limitations. First, we show that trait-specific selections of eigenvectors often diverge, sometimes producing inconsistent results, and that using their union offers stronger control of phylogenetic non-independence. Second, we evaluate robust regression estimators within PVR, demonstrating that PVR-MM - and in most cases PVR-L2, the standard OLS estimator - maintains high accuracy under non-stationary evolutionary shifts where other non-robust methods fail. Third, through simulation, we reassess the SPV [≥] 10 rule, showing that PVR tolerates eigenvector counts well beyond this threshold, offering greater flexibility while requiring attention to potential overfitting. Extensive simulations across diverse trees and evolutionary scenarios confirm that the optimized framework improves accuracy and robustness. By addressing key aspects of eigenvector selection, regression, and model complexity, our findings strengthen the reliability and applicability of PVR.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Hidden Side Of Diversity: Effects Of Imperfect Detection On Multiple Dimensions Of Biodiversity 91%
- Antecedent effect models as an exploratory tool to link climate drivers to herbaceous perennial population dynamics data 90%
- SAMPLE: an R package to estimate sampling effort for species' occurrence rates. 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.