Reference-based variant detection with varseek
Rich, J. M.; Luebbert, L.; Sullivan, D. K.; Rosa, R.; Pachter, L.
Show abstract
Variant detection from sequencing data is fundamental for genomics and is the first step in a wide range of applications, ranging from genome-wide association studies to disease diagnosis. Widely used tools for variant detection utilize a de novo approach that is based on a combination of read mapping algorithms and statistical methods for identifying genetic variation from error-prone sequencing data. This approach has been successful, although the detection of insertion and deletion variants, as well as the detection of variants from low-coverage data, remain challenging problems. We introduce varseek, a reference-based approach to variant detection that provides large improvements in performance in these challenging cases. The varseek approach utilizes a k-mer pseudoalignment approach, which provides the ability to identify variants at single-cell resolution in single-cell transcriptomics data. We showcase the versatility and performance of varseek for detecting tumor-specific COSMIC variants in glioblastoma single-cell sequencing.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Long-read Individual-molecule Sequencing Reveals CRISPR-induced Genetic Heterogeneity in Human ESCs 96%
- Accurate variant effect estimation in FACS-based deep mutational scanning data with Lilace 96%
- SlideCNA: Spatial copy number alteration detection from Slide-seq-like spatial transcriptomics data 96%
Similar papers in this journal
Similar papers in this journal
- Comprehensive benchmarking of methods for mutation calling in circulating tumor DNA 96%
- Identity-by-descent detection across 487,409 British samples reveals fine-scale population structure, evolutionary history, and trait associations 95%
- ProSolo: Accurate Variant Calling from Single Cell DNA Sequencing Data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.