Back

PhylUp: phylogenetic alignment building with custom taxon sampling

Kandziora, M.

2020-12-22 bioinformatics
10.1101/2020.12.21.394551 bioRxiv
Show abstract

In recent years it has become easier to reconstruct large-scale phylogenies with more or less automated workflows. However, they do not permit to adapt the taxon sampling strategy for the clade of interest. While most tools permit a single representative per taxon, PhylUp - the workflow presented here - enables to use different sampling strategies for different taxonomic ranks, as often needed for molecular dating analyses or for a large outgroup sampling. While PhylUp focuses on user-defined sampling strategies, it also facilitates the updating of alignments with new sequences from local and online sequence databases and their incorporation into existing alignments. To start a PhylUp run at least one sequence per locus has to be provided, PhylUp then adds new sequences to the existing one by internally using BLAST to find similar sequences and filters them according to user settings. Taxonomic sampling is increased compared to available tools and the custom taxonomic sampling allows to use automated workflows for new research fields. The workflow is presented in detail and I demonstrate the usability.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.