Back

Phyling: phylogenetic inference from annotated genomes

Tsai, C.-H.; Stajich, J. E.

2025-08-01 bioinformatics
10.1101/2025.07.30.666921 bioRxiv
Show abstract

Phyling is a fast, scalable, and user-friendly tool supporting phylogenomic reconstruction of species phylogenies directly from protein-encoded genomic data. It identifies orthologous genes by searching a samples protein sequences against a Hidden Markov Models marker set, containing single-copy orthologs, retrieved from the BUSCO database. In the final step, users can choose between consensus and concatenation strategies to construct the species tree from the aligned orthologs. Phyling efficiently resolves large phylogenies by optimizing memory usage and data processing. Its checkpoint system enables users to incrementally add or remove samples without repeating the entire search process. For analyses involving closely related taxa, Phyling supports the use of nucleotide coding sequences, which may capture phylogenetic signals missed by protein sequences. The benchmark results show that Phyling substantially runs faster than OrthoFinder, a Reciprocal Best Hit based method, while achieving equal or better accuracy.

Published in G3: Genes, Genomes, Genetics (predicted rank #14) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.