Back

Exploiting Large Datasets Improves Accuracy Estimation for Multiple Sequence Alignment

Cedillo, L.; Ruiz, H. R.; DeBlasio, D.

2022-05-24 bioinformatics
10.1101/2022.05.22.493004 bioRxiv
Show abstract

Multiple sequence alignment plays an important role in many important analyses. However, aligning multiple biological sequences is a complex task, thus many tools have been developed to align sequences under a biologically-inspired objective function. But these tools require a user-defined parameter vector, which if chosen incorrectly, can greatly impact down-stream analysis. Parameter Advising addresses this challenge of selecting input-specific parameter vectors by comparing alignments produced by a carefully constructed set of parameter configurations. In an ideal scenario, we would rank alignments based on their accuracy. However, in practice, we do not have a reference from which to calculate accuracy. Therefore, it becomes necessary to estimate the accuracy to rank the alignments. One solution involves the use of estimators such as Facet. The accuracy estimator Facet computes an estimate of accuracy as a linear combination of efficiently-computable feature functions. In this work we introduce two new estimators called Lead (short for Learned accuracy estimator from large datasets) which use the same underlying feature functions as Facet but are built on top of highly efficient machine learning protocols, allowing us to take advantage of a larger training corpus. Note about previous versionsA previous version of this paper was released on bioRxiv and presented the results of our previous study (Facet) with an error. This error has been corrected, and the conclusions made have been updated based on this new data. This corrected version stands as reference for anyone who may have encountered the versions with inaccuracies.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.