ModDotPlot - Rapid and interactive visualization of complex repeats
Sweeten, A. P.; Schatz, M. C.; Phillippy, A. M.
Show abstract
MotivationA common method for analyzing genomic repeats is to produce a sequence similarity matrix visualized via a dot plot. Innovative approaches such as StainedGlass have improved upon this classic visualization by rendering dot plots as a heatmap of sequence identity, enabling researchers to better visualize multi-megabase tandem repeat arrays within centromeres and other heterochromatic regions of the genome. However, computing the similarity estimates for heatmaps requires high computational overhead and can suffer from decreasing accuracy. ResultsIn this work we introduce ModDotPlot, an interactive and alignment-free dot plot viewer. By approximating average nucleotide identity via a k-mer-based containment index, ModDotPlot produces accurate plots orders of magnitude faster than StainedGlass. We accomplish this through the use of a hierarchical modimizer scheme that can visualize the full 128 Mbp genome of Arabidopsis thaliana in under 5 minutes on a laptop. ModDotPlot is bundled with a graphical user interface supporting real-time interactive navigation of entire chromosomes. Availability and ImplementationModDotPlot is available at https://github.com/marbl/ModDotPlot. Contactalex.sweeten@nih.gov, adam.phillippy@nih.gov
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MapGL: Inferring evolutionary gain and loss of short genomic sequence features by phylogenetic maximum parsimony. 94%
- SpectralTAD: an R package for defining a hierarchy of Topologically Associated Domains using spectral clustering 94%
- nPoRe: n-Polymer Realigner for improved pileup variant calling 94%
Similar papers in this journal
- Sensitive and error-tolerant annotation of protein-coding DNA with BATH 95%
- WAS IT A MATch I SAW? Approximate palindromes lead to overstated false match rates in benchmarks using reversed sequences 94%
- K2R: Tinted de Bruijn Graphs implementation for efficient read extraction from sequencing datasets 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.