Non-linear phylogenetic regression using regularized kernels
Rosas-Puchuri, U.; Santaquiteria, A.; Khanmohammadi, S.; Solis-Lemus, C.; Betancur-R., R.
Show abstract
O_LIPhylogenetic regression is a type of Generalized Least Squares (GLS) method that incorporates a covariance matrix based on the evolutionary relationships between species (i.e., phylogenetic relationships). While this method has found widespread use in hypothesis testing via comparative phylogenetic methods, such as phylogenetic ANOVA, its ability to account for non-linear relationships has received little attention. C_LIO_LITo address this issue, we utilized GLS in a high-dimensional feature space, employing linear combinations of transformed data to account for non-linearity, a common approach in kernel regression. We analyzed two biological datasets using both Radial Basis Function (RBF) and linear kernel transformations. The first dataset contained morphometric data, while the second dataset comprised discrete trait data and diversification rates as labels. Hyperparameter tuning of the model was achieved through cross-validation rounds in the training set. C_LIO_LIIn the tested biological datasets, regularized kernels reduced the error rate (as measured by RMSE) by around 20% compared to linear-based regression when data did not exhibit linear relationships. In simulated datasets, the error rate decreased almost exponentially with the level of non-linearity. C_LIO_LIThese results show that introducing kernels into phylogenetic regression analysis presents a novel and promising tool for complementing phylogenetic comparative methods. We have integrated this method into Python package named phyloKRR, which is freely available at: https://github.com/Ulises-Rosas/phylokrr. C_LI
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.