Skiver: Alignment-free Estimation of Sequencing Error Rates and Spectra using (k, v)-mer Sketches
Gu, Z.; Sharma, P.; Wong, L.; Nagarajan, N.
Show abstract
BackgroundQuality control of sequencing datasets is an important first step in numerous bioinformatics pipelines such as mapping, variant calling, and assembly. Existing methods typically rely on alignment results or quality scores. However, the reference genome is not always available for mapping, and uncalibrated quality scores may yield biased estimates of error rates. ResultsWe present skiver, a reference-free and alignment-free framework that estimates sequencing errors using (k, v)-mer sketches. By identifying the consensus through the sketched (k, v)-mers, skiver estimates survival and hazard rates that capture positional information of sequencing errors. Across simulated and real datasets from various sequencing platforms, skiver accurately recovers error rates and spectra. It also reliably handles complex datasets containing multiple strains, alleles, and repetitive regions through an outlier filtering strategy. Skiver is computationally efficient and provides a lightweight solution for error profiling in high-throughput sequencing. Availability and ImplementationThe implementation of skiver is available at https://github.com/GZHoffie/skiver, and the dataset and scripts for reproducibility are available at https://github.com/GZHoffie/skiver-test.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Accurate Estimation of Molecular Counts from Amplicon Sequence Data with Unique Molecular Identifiers 96%
- Demixer: A probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample 96%
- SECEDO: SNV-based subclone detection using ultra-low coverage single-cell DNA sequencing 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.