Using readmers and hapmers in assessing phase switching after read error correction of Oxford Nanopore Sequences
Elbers, J. P.; Horner, D.; Loewenstern, T.; Laccone, F.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWMethods for sequence error correction can improve sequence accuracy; however, there can be unintended errors added during error correction. One such example is phase switching, whereby sequences derived from genomes containing more than one parental copy have contributions from more than one parental haplotype. Such switches are mistakes that can confound downstream analyses especially de novo genome assembly. While DNA sequences do not possess words such as linguistic languages, one can partition a DNA sequence into word-like objects called k-mers. K-mers include pieces of DNA sequence of k length that can describe various properties of DNA sequences. With regard to phase switching, there are k-mers present in one parental haplotype not found in other(s). These so-called hapmers can represent inaccuracies in that parental haplotypes assembly but also correct, unique DNA variation. Here we investigated the effect of DNA sequence error correction on phase switching at the sequence/read level. Using several error-correction methods, we find all methods tested are similar to raw, presumably, phase-switch-free Oxford Nanopore Technologies (ONT) sequences in the percentage of readmers (k-mers from the ONT sequences) matching one parental haplotypes hapmers. This work demonstrates an efficient method to assess if an error-correction method has introduced phase switching implemented in the Julia programming language.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads 96%
- Whole-genome long-read sequencing downsampling and its effect on variant calling precision and recall 96%
- Haplocheck: Phylogeny-based Contamination Detection in Mitochondrial and Whole-Genome Sequencing Studies 95%
Similar papers in this journal
- Systematic benchmark of state-of-the-art variant calling pipelines identifies major factors affecting accuracy of coding sequence variant discovery 95%
- RecView: an interactive R application for viewing and locating recombination positions using pedigree data 95%
- Illuminating the dark side of the human transcriptome with TAMA Iso-Seq analysis 94%
Similar papers in this journal
- Concerning the eXclusion in human genomics: The choice of sex chromosome representation in the human genome drastically affects number of identified variants 95%
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 95%
- PAQman: reference-free ensemble evaluation of long-read eukaryotic genome assemblies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.