FastSpeciesTree: Fast and Scalable Species tree Inference
Holmes, J.; Kelly, S.
Show abstract
The generation of species trees from genomic data is an important step in understanding evolutionary relationships across the tree of life. With the increasing availability of genomic data, species tree inference methods which can easily and rapidly produce accurate species trees from large datasets are required to enable downstream analyses of these resources. Here we present FastSpeciesTree, a fully automated species tree inference method designed to leverage advances in rapid pairwise sequence alignment and scalable phylogenetic inference methods. We show that FastSpeciesTree is able to produce species trees with accuracies equivalent to trees produced through manual-curation and best-practice approaches. We further demonstrate that FastSpeciesTree can produce these high accuracy phylogenies substantially faster than any competitor method. Finally, we show that FastSpeciesTree is capable of building a species tree comprising almost the entirety of the RefSeq databases available proteomes (n = 33285) in less than 12 hours using standard computing resources. FastSpeciesTree is executed as a single command and can be found at https://github.com/OrthoFinder/FastSpeciesTree.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Rtapas: An R package to assess cophylogenetic signal between two evolutionary histories 94%
- OpenTree: A Python package for Accessing andAnalyzing data from the Open Tree of Life 94%
- Improved robustness to gene tree incompleteness, estimation errors, and systematic homology errors with weighted TREE-QMC 94%
Similar papers in this journal
- NGSEP 4: Efficient and Accurate Identification of Orthogroups and Whole Genome Alignment 95%
- SCRAPP: A tool to assess the diversity of microbial samples from phylogenetic placements 95%
- TREEasy: an automated workflow to infer gene trees, species trees, and phylogenetic networks from multilocus data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.