HamHeat: A fast and simple package for calculating Hamming distance from multiple sequence data for heatmap visualization
Rakov, A. V.; Schifferli, D. M.; Liu, S.-L.; Mastriani, E.
Show abstract
The problem of fast calculation of Hamming distance inferred from many sequence datasets is still not a trivial task. Here, we present HamHeat, as a new software package to efficiently calculate Hamming distance for hundreds of aligned protein or DNA sequences of a large number of residues or nucleotides, respectively. HamHeat uses a unique algorithm with many advantages, including its ease of use and the execution of fast runs for large amounts of data. The package consists of three consecutive modules. In the first module, the software ranks the sequences from the most to the least frequent variant. The second module uses the most common variant as the reference sequence to calculate the Hamming distance of each additional sequence based on the number of residue or nucleotide changes. A final module formats all the results in a comprehensive table that displays the sequence ranks and Hamming distances. Availability and implementationHamHeat is based on Python 3 and AWK, runs on Linux system and is available under the MIT License at: https://github.com/alexeyrakov/HamHeat. Contactrakovalexey@gmail.com
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MLDSP-GUI: An alignment-free standalone tool with an interactive graphical user interface for DNA sequence comparison and analysis 96%
- BamToCov: an efficient toolkit for sequence coverage calculations 95%
- ganon: precise metagenomics classification against large and up-to-date sets of reference sequences 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.