Iterative deletion of gene trees detects extreme biases in distance-based phylogenomic coalescent analyses
Gatesy, J.; Sloan, D. B.; Warren, J. M.; Simmons, M. P.; Springer, M. S.
Show abstract
Summary coalescent methods offer an alternative to the concatenation (supermatrix) approach for inferring phylogenetic relationships from genome-scale datasets. Given huge datasets, broad congruence between contrasting phylogenomic paradigms is often obtained, but empirical studies commonly show some well supported conflicts between concatenation and coalescence results and also between species trees estimated from alternative coalescent methods. Partitioned support indices can help arbitrate these discrepancies by pinpointing outlier loci that are unjustifiably influential at conflicting nodes. Partitioned coalescence support (PCS) recently was developed for summary coalescent methods, such as ASTRAL and MP-EST, that use the summed fits of individual gene trees to estimate the species tree. However, PCS cannot be implemented when distance-based coalescent methods (e.g., STAR, NJst, ASTRID, STEAC) are applied. Here, this deficiency is addressed by automating computation of partitioned coalescent branch length (PCBL), a novel index that uses iterative removal of individual gene trees to assess the impact of each gene on every clade in a distance-based coalescent tree. Reanalyses of five phylogenomic datasets show that PCBL for STAR and NJst trees helps quantify the overall stability/instability of clades and clarifies disagreements with results from optimality-based coalescent analyses. PCBL scores reveal severe missing taxa, apical nesting, misrooting, and basal dragdown biases. Contrived examples demonstrate the gross overweighting of outlier gene trees that drives these biases. Because of interrelated biases revealed by PCBL scores, caution should be exercised when using STAR and NJst, in particular when many taxa are analyzed, missing data are non-randomly distributed, and widespread gene-tree reconstruction error is suspected. Similar biases in the optimality-based coalescent method MP-EST indicate that congruence among species trees estimated via STAR, NJst, and MP-EST should not be interpreted as independent corroboration for phylogenetic relationships. Such agreements among methods instead might be due to the common defects of all three summary coalescent methods.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Perfect Storm: Gene Tree Estimation Error, Incomplete Lineage Sorting, and Ancient Gene Flow Explain the Most Recalcitrant Ancient Angiosperm Clade, Malpighiales 98%
- Whole-genomes illuminate the drivers of gene tree discordance and the tempo of tinamou diversification (Aves: Tinamidae) 96%
- Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC 96%
Similar papers in this journal
- Discovering Fragile Clades And Causal Sequences In Phylogenomics By Evolutionary Sparse Learning 97%
- Major revisions in pancrustacean phylogeny with recommendations for resolving challenging nodes 96%
- Comprehensive species sampling and sophisticated algorithmic approaches refute the monophyly of Arachnida 96%
Similar papers in this journal
- Categorical edge-based analyses of phylogenomic data reveal conflicting signals for difficult relationships in the avian tree 97%
- Comprehensive taxon sampling and vetted fossils help clarify the time tree of shorebirds (Aves, Charadriiformes) 95%
- Peeling Back the Layers: First Phylogenomic Insights into the Ledebouriinae (Scilloideae, Asparagaceae) 95%
Similar papers in this journal
- Determining the probability of hemiplasy in the presence of incomplete lineage sorting and introgression 95%
- An estimate of the deepest branches of the tree of life from ancient vertically-evolving genes 94%
- Myoglobin primary structure reveals multiple convergent transitions to semi-aquatic life in the world smallest mammalian divers 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.