PhylUp: phylogenetic alignment building with custom taxon sampling
Kandziora, M.
Show abstract
In recent years it has become easier to reconstruct large-scale phylogenies with more or less automated workflows. However, they do not permit to adapt the taxon sampling strategy for the clade of interest. While most tools permit a single representative per taxon, PhylUp - the workflow presented here - enables to use different sampling strategies for different taxonomic ranks, as often needed for molecular dating analyses or for a large outgroup sampling. While PhylUp focuses on user-defined sampling strategies, it also facilitates the updating of alignments with new sequences from local and online sequence databases and their incorporation into existing alignments. To start a PhylUp run at least one sequence per locus has to be provided, PhylUp then adds new sequences to the existing one by internally using BLAST to find similar sequences and filters them according to user settings. Taxonomic sampling is increased compared to available tools and the custom taxonomic sampling allows to use automated workflows for new research fields. The workflow is presented in detail and I demonstrate the usability.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- PhyloMagnet: Fast and accurate screening of short-read meta-omics data using gene-centric phylogenetics 95%
- Treerecs: an integrated phylogenetic tool, from sequences to reconciliations 95%
- PoSeiDon: a Nextflow pipeline for the detection of evolutionary recombination events and positive selection 94%
Similar papers in this journal
Similar papers in this journal
- COInr and mkCOInr: Building and customizing a non-redundant barcoding reference database from BOLD and NCBI using a lightweight pipeline. 96%
- A snakemake toolkit for the batch assembly, annotation, and phylogenetic analysis of mitochondrial genomes and ribosomal genes from genome skims of museum collections. 95%
- debar, a sequence-by-sequence denoiser for COI-5P DNA barcode data 95%
Similar papers in this journal
- ClusTRace, a bioinformatic pipeline for analyzing clusters in virus phylogenies 94%
- The advantages and disadvantages of short- and long-read metagenomics to infer bacterial and eukaryotic community composition 94%
- HiTaC: a hierarchical taxonomic classifier for fungal ITS sequences compatible with QIIME2 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.