Back

Stack Mapping Anchor Points (SMAP): a versatile suite of tools for read-backed haplotyping

Schaumont, D.; Veeckman, E.; Van der Jeugt, F.; Haegeman, A.; van Glabeke, S.; Bawin, Y.; Lukasiewicz, J.; Blugeon, S.; Barre, P.; Leyva-Perez, M. d. l. O.; Byrne, S. L.; Dawyndt, P.; Ruttink, T.

2022-03-13 genomics
10.1101/2022.03.10.483555 bioRxiv
Show abstract

Here we present SMAP, a software package that implements a suite of computational tools to extract multi-allelic haplotypes using read-backed haplotyping. SMAP tools first perform accurate read processing and analyze read mapping distributions across sample sets. Then, two complementary modules can be invoked for haplotype calling: SMAP haplotype-sites combines known Single Nucleotide Polymorphisms (SNPs) and/or read mapping position polymorphisms (SMAPs) to reconstruct compressed, read-reference-encoded haplotype strings. In contrast, SMAP haplotype-window works independent of prior knowledge of polymorphisms, groups reads by locus, defines a window enclosed between two custom border sequences, and retains the entire corresponding DNA sequence as haplotype. Haplotype-window is, among many applications, especially useful for high-throughput CRISPR/Cas mutation screens. Either way, SMAP creates a single integrated haplotype call table across all loci and samples. SMAP haplotyping is extremely versatile and can be applied to highly multiplex amplicon sequencing (HiPlex), Shotgun (e.g. whole genome shotgun (WGS) sequencing, probe capture and RNA-Seq), or Genotyping-by-Sequencing (GBS) data; and to Illumina short reads, PacBio and MinION long reads. SMAP creates discrete genotype calls for individuals of any ploidy or quantitative haplotype frequency spectra for Pool-Seq data, and can scale from tens to thousands of loci and/or samples. SMAP, including the source code written in Python is available at https://gitlab.com/truttink/smap, and a detailed user manual and guidelines for accurate read processing is available at https://ngs-smap.readthedocs.io/, under the GNU Affero General Public License v3.0.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.