Back

Clumps: A Sequential Clustering Approach to Partitioning Sets of Phylogenies with Non-Identical Leaf Sets

Serra Silva, A.; Wilkinson, M.

2025-03-19 bioinformatics
10.1101/2025.03.18.641816 bioRxiv
Show abstract

Post-processing of sets of inferred phylogenetic trees often focuses on the canonical case of consensus of (multi)sets of phylogenetic trees on the same leaf set. However, with growing numbers of phylogenomic studies resorting to summary and/or supertree methods to obtain a phylogeny, the amount of (multi)sets of trees with non-identical leaf sets also increases. In an attempt identify trees with non-identical leaf sets that are topologically similar, we define a new sequential subsetting approach, "clumps of trees", based on the distance between any tree in a set and the sets supertree. While clumps were developed with trees with non-identical leaf sets in mind, they can be applied to (multi)sets of trees with identical taxonomic sampling. Unlike islands of trees, clumps will not always be mutually exclusive, thus making them more similar to families of trees.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.