Accurate Multiple Sequence Alignment of Ultramassive Genome Sets
Lai, X.; Luan, H.; Tian, P.
Show abstract
With ever increasing sequencing efficiency, there is a pressing need to tackle a presently intractable task of accurate multiple sequence alignment (MSA) for ultramassive genome sets of millions and beyond. Additionally, efficient graph and probabilistic representations for downstream analysis are in dire lack. In light of these challenges, we develop a set of essentially linearly scalable algorithms, including that for constructing directed acyclic graphs, for training tiled profile hidden Markov models and for conducting alignment on such graphs. The power of these algorithms is demonstrated by both significantly improved accuracy and tremendous acceleration of SAR-CoV-2 MSA by three observed to five projected orders of magnitude for genome set sizes ranging from 40,000 to 4 million when compared with widely utilized MAFFT. Future application to other viral species and extension to more complex genomes will prove this algorithm set as a cornerstone for the coming era of ultramassive genome sets.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Adversarial domain translation networks for fast and accurate integration of large-scale atlas-level single-cell datasets 94%
- Improving the prediction of protein stability changes upon mutations by geometric learning and a pre-training strategy 93%
- An error correction strategy for image reconstruction by DNA sequencing microscopy 93%
Similar papers in this journal
- PAST: latent feature extraction with a Prior-based self-Attention framework for Spatial Transcriptomics 94%
- Cross-species cell-type assignment of single-cell RNA-seq by a heterogeneous graph neural network 94%
- CoRAL accurately resolves extrachromosomal DNA genomestructures with long-read sequencing 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.