Third generation indexing for third generation sequencing
Jighly, A.
Show abstract
Indexing of DNA sequences is the art of sorting massive genomic data in a user-friendly structure to enable rapid accessing and comparing of different patterns in the data. Current genome assemblers use general algorithms for string indexing that do not exploit the special structural arrangement of genomes. Here, I am proposing a new algorithm that indexes only the configuration of microsatellite motifs along reads assuming that the order of microsatellites will be the same in overlapped sequences. The index size is >1000 times smaller than currently used indices and it has higher tolerance to the high error rates produced by third generation sequencing platforms. The results showed that the proposed algorithm can rapidly detect overlaps among considerable proportion of uncorrected long reads (~50% of all simulated base pairs with average read size of 8.16 kb and total error rates of 14.4%) to build large initial contigs. Unassembled reads can be then mapped to these contigs or can be assembled with them with currently used algorithms. Thus, the proposed algorithm can efficiently be used as an initial stage to significantly reduce the number of pairwise sequence comparisons among reads and/or references and improve the performance of different software but not replacing them. The algorithm was also useful for comparative genomics and detect large locally colinear blocks and structural variations among ten saccharomyces cerevisiae strains. The proposed algorithm has the power to make de novo assembly of individuals as routine activity which can lead to more accurate variant calling and pan genomics.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Dotplotic: a lightweight visualization tool for BLAST+ alignments and genomic annotations 96%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 94%
- USAT: a Bioinformatic Toolkit to Facilitate Interpretation and Comparative Visualization of Tandem Repeat Sequences 94%
Similar papers in this journal
- BC-store: a program for mgiseq barcode sets analysis 94%
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 94%
- metaVaR: introducing metavariant species models for reference-free metagenomic-based population genomics 94%
Similar papers in this journal
Similar papers in this journal
- Robust and efficient software for reference-free genomic diversity analysis of GBS data on diploid and polyploid species 95%
- An exploration of assembly strategies and quality metrics on the accuracy of the Knightia excelsa (rewarewa) genome. 95%
- Chromosome-level hybrid de novo genome assemblies as an attainable option for non-model organisms 94%
Similar papers in this journal
- High-fidelity (repeat) consensus sequences from short reads using combined read clustering and assembly 95%
- RecView: an interactive R application for viewing and locating recombination positions using pedigree data 95%
- Nanopore sequencing and comparative genome analysis confirm lager-brewing yeasts originated from a single hybridization 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.