Back

Best: A Tool for Characterizing Sequencing Errors

Liu, D.; Belyaeva, A.; Shafin, K.; Chang, P.-C.; Carroll, A.; Cook, D.

2022-12-23 bioinformatics
10.1101/2022.12.22.521488 bioRxiv
Show abstract

SummaryPlatform-dependent sequencing errors must be understood to develop accurate sequencing technologies. We propose a new tool, best (Bam Error Stats Tool), for efficiently quantifying and summarizing error types in sequenced reads. best ingests reads aligned to a high-quality reference assembly and produces per-read metrics, summary statistics, and stratified metrics across genomic intervals. We show that best is 16 times faster than a prior method. In addition to being useful to support development that improves the accuracy of sequencing platforms, best can also be applied to evaluate and improve other experimental factors such as library preparation and error correction methods. Availability and implementationbest is an open-source command-line utility available on Github (github.com/google/best) under an MIT license. Contactdanielecook@google.com

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.