reconcILS: A gene tree-species tree reconciliation algorithm that allows for incomplete lineage sorting
Mishra, S.; Smith, M. L.; Hahn, M. W.
Show abstract
Reconciliation algorithms infer the evolutionary history of individual gene trees given a species tree. Many reconciliation algorithms consider only duplication and loss events (and sometimes horizontal transfer), ignoring effects of the coalescent process, including incomplete lineage sorting (ILS). Here, we present a new heuristic algorithm for carrying out reconciliation that accurately accounts for ILS by treating it as a series of nearest neighbor interchange (NNI) events. For discordant branches of the gene tree identified by last common ancestor (LCA) mapping, our algorithm recursively chooses the optimal history by comparing the cost of duplication and loss to the cost of NNI and loss. We demonstrate the accuracy of our new method, which we call reconcILS, using a new simulation engine (dupcoal) that generates gene trees produced by the interaction of duplication, loss, and ILS under the MSC-DL model. Despite being a heuristic method, reconcILS is much more accurate than models that ignore ILS, and at least as accurate or better than leading methods that can model ILS, while also able to handle much larger datasets. We demonstrate the use of reconcILS by applying it to a dataset of 23 primate genomes, highlighting its accuracy compared to standard methods in the presence of large amounts of ILS.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- PhyloCoalSimulations: A simulator for network multispecies coalescent models, including a new extension for the inheritance of gene flow 97%
- Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC 97%
- Estimating Waiting Distances Between Genealogy Changes under a Multi-Species Extension of the Sequentially Markov Coalescent 97%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.