NanoSquiggleVar: A method for direct analysis of targeted variants based on nanopore sequencing signals
Lang, J.
Show abstract
BackgroundNanopore sequencing is a fourth-generation sequencing technology that has developed rapidly in recent years. It has long sequencing read lengths and does not require the polymerase chain reaction to be performed. These characteristics give it unique advantages over the next-generation sequencing technology under certain usage scenarios. The number of bioinformatics analysis algorithms and/or tools developed with nanopore sequencing has increased sharply during the past years, undoubtedly providing great help and support for the application of nanopore sequencing in scientific research and practical scenarios. ResultsWe developed NanoSquiggleVar, a method for direct analysis of targeted variants based on nanopore sequencing signals. It first establishes a set of wild-type and mutant-type target signals within the same experimental and sequencing system, named wild squiggle set and variant squiggle set, respectively. In each sequencing iteration, the signal is sliced into fragments by a moving window of 1-unit step size. Then, dynamic time warping is used to compare the signal squiggles to the detected variants. Point mutations, insertions and deletions (indels), and homopolymer sequences were simulated and generated by Scrappie and then analyzed and evaluated with NanoSquiggleVar. We found that all of these variants were efficiently detected and discriminated, and the results were consistent with the expectations. ConclusionsNanoSquiggleVar can directly identify targeted variants from the nanopore sequencing electrical signal without the requirement of base calling, sequence alignment, or variant detection with downstream analysis. We hope that this method can complement targeted variant detection using nanopore sequencing and potentially serve as a reference for real-time sequencing and analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An ultra-sensitive T-cell receptor detection method for TCR-Seq and RNA-Seq data 94%
- Identify phage hosts from metaviromic short reads based on deep learning and Markov chain model 93%
- Searching and mapping genomic subsequences in nanopore raw signals through novel dynamic time warping algorithms 93%
Similar papers in this journal
- 3rd-ChimeraMiner: A pipeline for integrated analysis of whole genome amplification generated chimeric sequences using long-read sequencing 94%
- DISMIR: a deep learning-based cancer-detection method by integrating DNA sequence and methylation information of individual cell-free DNA reads 94%
- Choice of assemblers has a critical impact on de novo assembly of SARS-CoV-2 genome and characterizing variants 93%
Similar papers in this journal
- Giraffe: a tool for comprehensive processing and visualization of multiple long-read sequencing data 95%
- UMI-Gen: a UMI-based reads simulator for variant calling evaluation in paired-end sequencing NGS libraries 93%
- LongReadSum: A fast and flexible quality control and signal summarization tool for long-read sequencing data 93%
Similar papers in this journal
- PtWAVE: A High-Sensitive deconvolution software of sequencing trace for the Detection of Large Indels in Genome Editing 95%
- gencore: an efficient tool to generate consensus reads for error suppressing and duplicate removing of NGS data 95%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 94%
Similar papers in this journal
- Genomic style: yet another deep-learning approach to characterize bacterial genome sequences 94%
- ClusterV-Web: A User-Friendly Tool for Profiling HIV Quasispecies and Generating Drug Resistance Reports from Nanopore Long-Read Data 92%
- DANGER analysis: Risk-averse on/off-target assessment for CRISPR editing without a reference genome 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.