Back

Accurate Multiple Sequence Alignment of Ultramassive Genome Sets

Lai, X.; Luan, H.; Tian, P.

2024-09-23 bioinformatics
10.1101/2024.09.22.613454 bioRxiv
Show abstract

With ever increasing sequencing efficiency, there is a pressing need to tackle a presently intractable task of accurate multiple sequence alignment (MSA) for ultramassive genome sets of millions and beyond. Additionally, efficient graph and probabilistic representations for downstream analysis are in dire lack. In light of these challenges, we develop a set of essentially linearly scalable algorithms, including that for constructing directed acyclic graphs, for training tiled profile hidden Markov models and for conducting alignment on such graphs. The power of these algorithms is demonstrated by both significantly improved accuracy and tremendous acceleration of SAR-CoV-2 MSA by three observed to five projected orders of magnitude for genome set sizes ranging from 40,000 to 4 million when compared with widely utilized MAFFT. Future application to other viral species and extension to more complex genomes will prove this algorithm set as a cornerstone for the coming era of ultramassive genome sets.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.