Dromi: Python package for parallel computation of similarity measures among vector-encoded sequences
Sanz Moreta, L.
Show abstract
Calculating similarities among sequences (i.e biological sequences) can be a challenging task. Here I introduce Dromi, a simple python package that can compute different similarity measurements (i.e percent identity, cosine similarity, kmer similarities) across aligned vector-encoded sequences. This is a crucial step required to perform both upstream and downstream sequence machine learning tasks such as sequence clustering [1, 2, 3], sequence analysis [4] and other pre- or post-processing demands on sequences. Additionally, this package introduces the calculation of the measure referred as positional weights. These represent the cosine similarities or residue-conservation across sequence elements (i.e amino acids in peptide sequences) in the same site (column). The program can also deal with sequences of variable length since end-padded positions are not considered for the calculations. The presented implementations are an incorporation into the arsenal of tools to measure similarity among small peptide sequences such as epitopes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Constructing benchmark test sets for biological sequence analysis using independent set algorithms 94%
- MCell4 with BioNetGen: A Monte Carlo Simulator of Rule-Based Reaction-Diffusion Systems with Python Interface 94%
- gmxapi: a Gromacs-native Python interface for molecular dynamics with ensemble and plugin support 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.