Back

KoT: an automatic implementation of the K/θ method for species delimitation

Spöri, Y.; Stoch, F.; Dellicour, S.; Birky, C. W.; Flot, J.-F.

2021-08-18 evolutionary biology
10.1101/2021.08.17.454531 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWK/{theta} is a method to delineate species that rests on the calculation of the ratio between the average distance K separating two putative species-level clades and the genetic diversity{theta} of these clades. Although this method is explicitly rooted in population genetic theory, it was never benchmarked due to the absence of a program allowing automated analyses. For the same reason, its application by hand was limited to small datasets of a few tens of sequences. We present an automatic implementation of the K/{theta} method, dubbed KoT (short for "K over Theta"), that takes as input a FASTA file, builds a neighbour-joining tree, and returns putative species boundaries based on a user-specified K/{theta} threshold. This automatic implementation avoids errors and makes it possible to apply the method to datasets comprising many sequences, as well as to test easily the impact of choosing different K/{theta} threshold ratios. KoT is implemented in Haxe, with a javascript webserver interface freely available at https://eeg-ebe.github.io/KoT/

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.