Unifying phylogenetic traversal and deep learning to guide tree exploration
Collienne, L.; Richman, H.; Rich, D. H.; Barker, M.; Jennings-Shaffer, C.; Matsen, F. A.
Show abstract
Deep learning offers hope for more efficient phylogenetic inference methods. However, it has yet to have the transformative effect on phylogenetics that it has had in other fields. Here we present a novel approach that combines deep learning with concepts behind current successful phylogenetic algorithms. Specifically, we give the deep learning algorithm access to the output of a phylogenetic dynamic program on the sequence alignment, rather than the raw sequence alignment. The algorithm then learns features based on these phylogenetically processed versions of the sequence data, which provides information that could inform local tree search. For this paper, our goal is simple: predict for each edge in a tree whether it is in a maximum parsimony tree or not. Our model consists of a recurrent neural network that learns features while traversing the input tree, which are used to classify the edge. The model makes high-quality predictions for this NP-complete problem on simulated and empirical datasets for trees of various sizes, and we believe is a stepping stone towards efficient phylogenetic inference using deep learning.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- PhyloCNN: Improving tree representation and neural network architecture for deep learning from trees in phylodynamics and diversification studies 96%
- Deep learning and likelihood approaches for viral phylogeography converge on the same answers whether the inference model is right or wrong 95%
- ConvexML: Fast and accurate branch length estimation under irreversible mutation models, illustrated through applications to CRISPR/Cas9-based lineage tracing 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.