Poplar: A Phylogenetics Pipeline
Krishnakumar, R.; Koning, E.
Show abstract
MotivationGenerating phylogenetic trees from genomic data is essential in understanding biological systems. Each step of this complex process has received extensive attention in the literature, and has been significantly streamlined over the years. Given the volume of publicly available genetic data, obtaining genomes for a wide selection of known species is straightforward. However, analyzing that same data in order to generate a phylogenetic tree is a multi-step process with legitimate scientific and technical challenges, and often requires a significant input from a domain-area scientist. ResultsWe present Poplar, a new, streamlined computational pipeline, to address the computational logistical issues that arise when constructing phylogenetic trees. It provides a framework that runs state-of-the-art software for essential steps in the phylogenetic pipeline, beginning from a genome with or without an annotation, and resulting in a species tree. Running Poplar requires no external databases. In the execution, it enables parallelism for execution for clusters and cloud computing. The trees generated by Poplar match closely with state-of-the-art published trees. The usage and performance of Poplar is far simpler and quicker than manually running a phylogenetic pipeline. Availability and ImplementationFreely available on GitHub at https://github.com/sandialabs/poplar. Implemented using Python and supported on Linux. Supplementary InformationNewick versions of the reference and generated trees.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- PlantTribes2: tools for comparative gene family analysis in plant genomics 94%
- PubPlant: a continuously updated online resource for sequenced and published plant genomes 92%
- In silico evidence for the utility of parsimonious root phenotypes for improved vegetative growth and carbon sequestration under drought 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.