An improved mode of running PASTA
Yang, Q.; Warnow, T.
Show abstract
PASTA is a method for estimating alignments and trees that has been able to provide excellent accuracy on large sequence datasets. By design, PASTA operates using iteration, in which the tree from the previous iteration is used to inform a divide-and-conquer strategy during which a new alignment is computed on the sequence dataset, and then a new maximum likelihood tree is estimated on the new alignment. In its default setting, PASTA runs for three iterations and returns that alignment/tree pair from the last iteration. Here we use both biological and simulated nucleotide datasets to show that returning the alignment/tree pair that has the best maximum likelihood score improves on the default usage.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Evaluating probabilistic programming and fast variational Bayesian inference in phylogenetics 94%
- Fine-Tuning Protein Language Models Unlocks the Potential of Underrepresented Viral Proteomes 93%
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 93%
Similar papers in this journal
- Unspecific binding but specific disruption of the group I intron by the StpA chaperone 94%
- Evaluating DCA-based method performances for RNA contact prediction by a well-curated dataset 92%
- Unraveling Unbreakable Hairpins: Characterizing RNA secondary structures that are persistent after dinucleotide shuffling 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.