Characterizing and addressing error modes to improve sequencing accuracy
Kruglyak, S.; Altomare, A.; Ambroso, M.; Dien, V.; Lajoie, B. R.; Wiseman, K. N.; Levy, S. E.; Kellinger, M.
Show abstract
The accuracy of a sequencing platform has traditionally been measured by the %Q30, or percentage of data exceeding a basecall accuracy of 99.9%. Improvements to accuracy beyond Q30 may be beneficial for certain applications such as the identification of low frequency alleles or the improvement of reference genomes. Here we demonstrate how we achieved over 70% Q50 (99.999% accuracy) data on the AVITI sequencer. This level of accuracy required us to not only improve sequencing quality but also to mitigate library preparation errors and analysis artifacts.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Characterization and Mitigation of Fragmentation Enzyme-Induced Dual Stranded Artifacts 94%
- Fast and memory-efficient mapping of short bisulfite sequencing reads using a two-letter alphabet 93%
- Whole genome sequencing with AVITI and NovaSeq X Plus reveals comparable performance with contextual biases 93%
Similar papers in this journal
- Low-pass sequencing plus imputation using avidity sequencing displays comparable imputation accuracy to sequencing by synthesis while reducing duplicates 91%
- Optimized ChIP-exo for mammalian cells and patterned sequencing flow cells 91%
- WeavePop: A bioinformatics workflow to explore and analyze genomic variants of eukaryotic populations 90%
Similar papers in this journal
- TopoQual polishes circular consensus sequencing data and accurately predicts quality scores 95%
- DeepSelectNet: Deep Neural Network Based Selective Sequencing for Oxford Nanopore Sequencing 95%
- PIPETS: A statistically informed, gene-annotation agnostic analysis method to study bacterial termination using 3'-end sequencing. 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.