A New approximate matching compression algorithm for DNA sequences
Lazaro-Guevara, J. M.; Garrido, K. M.
Show abstract
1.Undeveloped countries like Guatemala, where access to high-speed internet connections is limited, downloading and sharing Biological information of thousands of Mega Bits is a huge problem for the beginning and development of Bioinformatics. Based on that information is an urgent necessity to find a better way to share this biological data. There is when the compression algorithms become relevant. With all this information in mind, born the idea of creating a new algorithm using redundancy and approximate selection. MethodsUsing the probability given by the transition matrix of the three-word tuple and relative frequencies. Calculating the relative and total frequencies given by the permutation formula (nr) and compressing 6 bits of information into 1 implementing the ASCII table code (0...255 characters, 28), using clusters of 102 DNA bases compacted into 17 string sequences. For decompressing, the inverse process must be done, except that the triplets must be selected randomly (or use a matrix dictionary, 4102). ConclusionThe compression algorithm has a better compression ratio than LZW and Huffmans algorithm. However, the time needed for decompressing makes this algorithm incompatible for massive data. The functionality as MD5sum need more research but is a promising helpful tool for DNA checking.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Artificial intelligence tool for the study of COVID-19 microdroplet spread across the human diameter and airborne space 95%
- Multi- Stage Feature Selection (MSFS) Algorithm for UWB- Based Early Breast Cancer Size Prediction 95%
- Prediction and control of COVID-19 infection based on a hybrid intelligent model 94%
Similar papers in this journal
- Extensive In Silico Analysis of the Functional and Structural Consequences of SNPs in Human ARX Gene 95%
- Viral miRNAs Confer Survival in Host Cells by Targeting Apoptosis Related Host Genes 93%
- Finding Consensus miRNAs Silencing KLF1 Expression as A Promising Therapeutic Option of Sickle Cell Anemia 93%
Similar papers in this journal
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 96%
- Principal Component Analysis applied directly to Sequence Matrix 94%
- Comparing protein-protein interaction networks of SARS-CoV-2 and (H1N1) influenza using topological features 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.