Improving long-read consensus sequencing accuracy with deep learning
Lal, A.; Brown, M.; Mohan, R.; Daw, J.; Drake, J.; Israeli, J.
Show abstract
The PacBio HiFi sequencing technology combines less accurate, multi-read passes from the same molecule (subreads) to yield consensus sequencing reads that are both long (averaging 10-25 kb) and highly accurate. However, these reads can retain residual sequencing error, predominantly insertions or deletions at homopolymeric regions. Here, we train deep learning models to polish HiFi reads by recognizing and correcting sequencing errors. We show that our models are effective at reducing these errors by 25-40% in HiFi reads from human as well as E. coli genomes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- HISAT-3N: a rapid and accurate three-nucleotide sequence aligner 96%
- Decoil: Reconstructing extrachromosomal DNA structural heterogeneity from long-read sequencing data 96%
- An Algorithm for Sequence Location Approximation using Nuclear Families (ASLAN) Validates Regions of the Telomere-to-Telomere Assembly and Identifies New Hotspots for Genetic Diversity 96%
Similar papers in this journal
- Lancet2: Improved and accelerated somatic variant calling with joint multi-sample local assembly graph 95%
- Needlestack: an ultra-sensitive variant caller for multi-sample next generation sequencing data 94%
- Scalable and efficient DNA sequencing analysis on different compute infrastructures aiding variant discovery 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.