Pangenome graph augmentation from unassembled long reads
Denti, L.; Bonizzoni, P.; Brejova, B.; Chikhi, R.; Krannich, T.; Vinar, T.; Hormozdiari, F.
Show abstract
Pangenomes are becoming increasingly popular data structures for genomics analyses due to their ability to compactly represent the genetic diversity within populations. Constructing a pangenome graph, however, is still a time-consuming and expensive process. A promising approach for pangenome construction consists of progressively augmenting a pangenome graph with additional high-quality assemblies. Currently, there is no method for augmenting a pangenome graph with unassembled reads from newly sequenced samples without first aligning the reads to a reference genome and performing variant calling and genotyping on the new individuals. In this work, we present the first assembly-free and mapping-free approach for augmenting an existing pangenome graph using unassembled long reads from an individual not already present in the pangenome. Our approach consists of finding sample specific sequences in reads using efficient indexes, clustering reads corresponding to the same novel variant(s), and then building a consensus sequence to be added to the pangenome graph for each variant separately. Using simulated reads based on Human Pangenome Reference Consortium (HPRC) assemblies, we demonstrate the effectiveness of the proposed approach for progressively augmenting the pangenome with long reads, without the need for de novo assembly or predicting genetic variants of the new sample. The software is freely available at https://github.com/ldenti/palss.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Freddie: Annotation-independent Detection and Discovery of Transcriptomic Alternative Splicing Isoforms 97%
- Pasa: Leverage population pangenome graph to scaffold prokaryote genome assemblies. 96%
- Accurate assembly of minority viral haplotypes from next-generation sequencing through efficient noise reduction 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.