Back

Repeat and haplotype aware error correction in nanopore sequencing reads with DeChat

Li, Y.; Chen, E.; Xu, J.; Zhang, W.; Zeng, X.; Liu, Y.; Luo, X.

2024-05-10 bioinformatics
10.1101/2024.05.09.593079 bioRxiv
Show abstract

Error self-correction is a pivotal first step in the analysis of long-read sequencing data. However, most existing methods for this purpose are primarily tailored for noisy sequencing data with error rates exceeding 5%, often collapsing true variants in repeats and haplotypes. Alternatively, some methods are heavily optimized for PacBio HiFi reads, leaving a gap in methods specifically designed for Nanopore R10 reads basecalled with high accuracy or super accuracy models, which typically have error rates below 2%. Here, we introduce DeChat, a novel approach specifically designed for Nanopore R10 reads. DeChat enables repeat- and haplotype-aware error correction, leveraging the strengths of both de Bruijn graphs and variant-aware multiple sequence alignment to create a synergistic approach. This approach avoids read overcorrection, ensuring that variants in repeats and haplotypes are preserved while sequencing errors are accurately corrected. Benchmarking experiments reveal that reads corrected using DeChat exhibit substantially fewer errors, ranging from several times to two orders of magnitude lower, compared to the current state-of-the-art approaches. Furthermore, the application of DeChat for error correction significantly improves genome assembly across various aspects. DeChat is implemented as a highly efficient, standalone, and user-friendly software and is publicly available at https://github.com/LuoGroup2023/DeChat.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.