Back

Optimized Phylogenetic Eigenvector Regression Provides Robust and Flexible Trait Correlation Analysis under Diverse Evolutionary Conditions

Chen, Z.-L.; Niu, D.-K.

2025-04-17 evolutionary biology
10.1101/2024.04.14.589420 bioRxiv
Show abstract

Phylogenetic eigenvector regression (PVR) is widely used in ecology and evolution by representing phylogenetic structure through separable eigenvectors. Despite this flexibility, its implementation faces three key challenges: (1) the selection of eigenvectors, (2) the reduced robustness of ordinary least-squares (OLS) regression under shift-like evolutionary heterogeneity, and (3) the applicability of conventional model complexity rules such as the "samples-per-variable (SPV) [≥] 10" guideline. Here, we propose an optimized PVR framework that addresses these limitations. First, we show that trait-specific selections of eigenvectors often diverge, sometimes producing inconsistent results, and that using their union offers stronger control of phylogenetic non-independence. Second, we evaluate robust regression estimators within PVR, demonstrating that PVR-MM - and in most cases PVR-L2, the standard OLS estimator - maintains high accuracy under non-stationary evolutionary shifts where other non-robust methods fail. Third, through simulation, we reassess the SPV [≥] 10 rule, showing that PVR tolerates eigenvector counts well beyond this threshold, offering greater flexibility while requiring attention to potential overfitting. Extensive simulations across diverse trees and evolutionary scenarios confirm that the optimized framework improves accuracy and robustness. By addressing key aspects of eigenvector selection, regression, and model complexity, our findings strengthen the reliability and applicability of PVR.

Published in Evolution (predicted rank #8) · training set

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.