Phyling: phylogenetic inference from annotated genomes
Tsai, C.-H.; Stajich, J. E.
Show abstract
Phyling is a fast, scalable, and user-friendly tool supporting phylogenomic reconstruction of species phylogenies directly from protein-encoded genomic data. It identifies orthologous genes by searching a samples protein sequences against a Hidden Markov Models marker set, containing single-copy orthologs, retrieved from the BUSCO database. In the final step, users can choose between consensus and concatenation strategies to construct the species tree from the aligned orthologs. Phyling efficiently resolves large phylogenies by optimizing memory usage and data processing. Its checkpoint system enables users to incrementally add or remove samples without repeating the entire search process. For analyses involving closely related taxa, Phyling supports the use of nucleotide coding sequences, which may capture phylogenetic signals missed by protein sequences. The benchmark results show that Phyling substantially runs faster than OrthoFinder, a Reciprocal Best Hit based method, while achieving equal or better accuracy.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Broccoli: combining phylogenetic and network analyses for orthology assignment 97%
- PhyloAln: a convenient reference-based tool to align sequences and high-throughput reads for phylogeny and evolution in the omic era 96%
- PhylteR: efficient identification of outlier sequences in phylogenomic datasets 95%
Similar papers in this journal
- NGSEP 4: Efficient and Accurate Identification of Orthogroups and Whole Genome Alignment 96%
- TREEasy: an automated workflow to infer gene trees, species trees, and phylogenetic networks from multilocus data 96%
- SCRAPP: A tool to assess the diversity of microbial samples from phylogenetic placements 95%
Similar papers in this journal
- ClusTRace, a bioinformatic pipeline for analyzing clusters in virus phylogenies 95%
- HiTaC: a hierarchical taxonomic classifier for fungal ITS sequences compatible with QIIME2 95%
- SANS ambages: phylogenomics with abundance-filter, multi-threading, and bootstrapping on amino-acid or genomic sequences 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.