A method for massively scalable phylogenetic network inference
Kolbow, N.; Kong, S.; Solis-Lemus, C.
Show abstract
Recent advancements in sequencing technologies have enabled large-scale phylogenomic analyses. While these analyses often rely on phylogenetic trees, increasing evidence suggests that non-treelike evolutionary events, such as hybridization and horizontal gene transfer, are prevalent in the evolutionary histories of many species, in which case tree-based models are insufficient. Phylogenetic networks can capture such complex evolutionary histories, but current methods for accurately inferring them lack scalability. Implicit network inference methods are fast but lack biological interpretability. Here, we introduce a novel method called InPhyNet that merges a set of non-overlapping, independently inferred level-1 networks into a unified topology, achieving linear scalability while maintaining high accuracy under the multispecies network coalescent model. We prove that a pipeline utilizing InPhyNet can be statistically consistent if the proper methodology is used. Using simulation, we infer networks with up to 200 taxa and show that divide-and-conquer pipelines utilizing InPhyNet allow for accurate network inference at scales and speeds previously unseen. Re-analyzing a phylogeny of 1,158 land plants with InPhyNet, we recover known reticulate events and illustrate how InPhyNet enables large-scale analyses of biologically meaningful reticulate phylogenies at previously unprecedented scales.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Inference of Phylogenetic Networks from Sequence Data using Composite Likelihood 97%
- DEPP: Deep Learning Enables Extending Species Trees using Single Genes 96%
- ConvexML: Fast and accurate branch length estimation under irreversible mutation models, illustrated through applications to CRISPR/Cas9-based lineage tracing 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.