TreeFormer: A transformer-based tree rearrangement operation for phylogenetic reconstruction
Ly-Trong, N.; Albert Matsen, F.; Minh, B. Q.
Show abstract
Phylogenetic inference is a fundamental problem in biology, which studies the origins and evolutionary relationships among species. Popular phylogenetic inference methods, such as IQ-TREE, RAxML, and PHYML, typically utilize heuristic tree search algorithms to seek a phylogenetic tree that maximizes the likelihood of the observed genetic data. However, tree search is time-consuming and often prone to local optima. To address these issues, we introduce TreeFormer, a new Transformer-based tree rearrangement operation for tree search. Experimental results show that TreeFormer achieves higher accuracy than FastTree 2 when reconstructing trees from real alignments with fewer than 1000 sites.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- GeneRax: A tool for species tree-aware maximum likelihood based gene tree inference under gene duplication, transfer, and loss. 98%
- Adaptive RAxML-NG: Accelerating Phylogenetic inference under Maximum Likelihood using dataset difficulty 97%
- AliSim: A Fast and Versatile Phylogenetic Sequence Simulator For the Genomic Era 97%
Similar papers in this journal
- AleRax: A tool for gene and species tree co-estimation and reconciliation under a probabilistic model of gene duplication, transfer and loss. 97%
- Build a Better Bootstrap and the RAWR Shall Beat a Random Path to Your Door: Phylogenetic Support Estimation Revisited 96%
- A Divide-and-Conquer Method for Scalable Phylogenetic Network Inference from Multi-locus Data 96%
Similar papers in this journal
- Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC 96%
- Inference of Phylogenetic Networks from Sequence Data using Composite Likelihood 96%
- PhyloFusion- Fast and easy fusion of rooted phylogenetic trees into rooted phylogenetic networks 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.