Back

Bcmap: fast alignment-free barcode mapping for linked-read sequencing data

Luepken, R.; Krannich, T.; Kehr, B.

2022-06-21 bioinformatics
10.1101/2022.06.20.496811 bioRxiv
Show abstract

MotivationThe bottleneck for genome analysis will soon shift from sequencing cost to computationally expensive read alignment. When only portions of the genome are of interest, read alignment on whole-genome sequencing (WGS) data can be replaced by a fast mapping step. ResultsWe propose an approach that can map both long reads or linked-read molecules, as a pre-processing step for downstream targeted analyses. Our "molecule mapping" approach, implemented in the tool MoleMap, uses a minimized open-addressing k-mer index of the reference genome and a fast k-mer clustering procedure. We demonstrate that MoleMaps accuracy is competitive with standard alignment and mapping tools and its running time outperforms other mapping tools on 32 threads by a factor of 3 to 8 and read alignment tools by a factor of 10 to 60. Its low memory footprint allows us to analyze whole genomes on a standard laptop computer. As proofs of concept, we use MoleMap to filter reads for local assembly of a known variant region that involves non-reference sequence and showcase its use in diagnosing a patient with a rare disease. Our work contributes to more scalable genome analysis and promotes WGS for targeted analyses. Availability and ImplementationSource code of MoleMap is available at https://github.com/kehrlab/molemap.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.