A Systematic Investigation of Overfitting in Maximum Likelihood Phylogenetic Inference
Togkousidis, A.; Gascuel, O.; Stamatakis, A.
Show abstract
Maximum Likelihood (ML) tree inference reconstructs phylogenies from Multiple Sequence Alignments (MSAs). Since MSAs are inherently noisy, ML tools may experience overfitting, whereby the inferred topology incorrectly models noise alongside the true phylogenetic signal. We statistically assess overfitting in ML tools using log-likelihood scores on unseen sites as primary metric. We deploy a 10-fold Monte Carlo cross-validation approach, partitioning 9,062 empirical and 6,342 simulated MSAs into training (80%) and testing (20%) sites. We conduct inferences using RAxML-NG, IQ-TREE, Fast-Tree, and RAxML-NG ES (a recently released Early Stopping version) on the training MSAs. We store all intermediate improved topologies and subsequently evaluate them on the testing sites. We perform a linear regression on the final segments (we use four distinct segment-window configurations) of the derived testing curves, and statistically evaluate the line slopes via the sign test. Our results indicate that ML tools do not overfit. For RAxML-NG (standard and ES) and IQ-TREE, the overall trend is non-significant for 86-98% of the empirical MSAs (across all four segment-window configurations), while for less than 1% the tools exhibit over-fitting. FastTree shows more positive trends, suggesting premature termination, especially on protein MSAs. Topological accuracy curves on simulated data further confirm that tools do not systematically diverge from the true topology. To complement these findings, we test whether strategies to mitigate overfitting can benefit ML inferences. To this end, we also benchmark a site-based holdout validation (HV) version of RAxML-NG. The results confirm that the overfitting is absent and also indicate that excluding MSA sites substantially reduces phylogenetic signal as well as accuracy.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Adaptive RAxML-NG: Accelerating Phylogenetic inference under Maximum Likelihood using dataset difficulty 98%
- From Easy to Hopeless - Predicting the Difficulty of Phylogenetic Analyses 98%
- GeneRax: A tool for species tree-aware maximum likelihood based gene tree inference under gene duplication, transfer, and loss. 97%
Similar papers in this journal
Similar papers in this journal
- Build a Better Bootstrap and the RAWR Shall Beat a Random Path to Your Door: Phylogenetic Support Estimation Revisited 97%
- AleRax: A tool for gene and species tree co-estimation and reconciliation under a probabilistic model of gene duplication, transfer and loss. 96%
- SODA: Multi-locus species delimitation using quartet frequencies 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.