Back

DIST: Distance-based Inference of Species Trees

de Jong, M. J.; Janke, A.

2025-05-08 bioinformatics
10.1101/2025.05.02.651899 bioRxiv
Show abstract

Inferring species trees from concatenated loci is often criticised for failing to account for gene tree discordance - particularly when using character-based methods. However, this criticism does not apply to distance-based concatenation trees, which can be shown to be statistically consistent even in anomaly zones. Building on this insight, we introduce DIST (Distance-based Inference of Species Trees), an intuitive and scalable method that infers species trees from population-level distance matrices containing multi-locus estimates of Dxy, FST or coalescence units ({tau}). DIST derives these values from between-individual sequence dissimilarity estimates, E(p), using basic equations from coalescence theory. Under certain conditions, DIST can also quantify gene tree discordance and distinguish whether it arises from gene flow or incomplete lineage sorting alone. While conceptually related to more sophisticated summary methods, DIST differs in that it does not seek the species tree which best explains a set of gene trees. Instead, it searches for the species tree which best explains an average gene tree, of which all branch lengths reflect mean coalescence time, E(t). Although this average gene tree is rarely observed empirically, it is approximated by an individual-level distance-based tree, traditionally referred to as a tree of individuals. The DIST algorithm is implemented in the R package SambaR, which now accepts input in the form of pairwise E(p) estimates.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.