Back

An improved mode of running PASTA

Yang, Q.; Warnow, T.

2020-09-01 bioinformatics
10.1101/2020.08.30.274217 bioRxiv
Show abstract

PASTA is a method for estimating alignments and trees that has been able to provide excellent accuracy on large sequence datasets. By design, PASTA operates using iteration, in which the tree from the previous iteration is used to inform a divide-and-conquer strategy during which a new alignment is computed on the sequence dataset, and then a new maximum likelihood tree is estimated on the new alignment. In its default setting, PASTA runs for three iterations and returns that alignment/tree pair from the last iteration. Here we use both biological and simulated nucleotide datasets to show that returning the alignment/tree pair that has the best maximum likelihood score improves on the default usage.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.