Development and Application of a Novel Simple Sequence Repeat Mining Algorithm Based on Regular Expression
Jia, Z.; Geng, R.; Wu, X.; Chen, S.; Tong, Y.; Yang, A.; Luo, C.; Ren, M.
Show abstract
Simple sequence repeats (SSRs) are molecular genetic markers that are powerful tools in genomics studies; SSR markers are routinely mined as a part of genetic workflows. Here, we developed a novel SSR mining algorithm based on regular expression that can reduce the complexity of commonly used SSR mining software. We used the following SSR mining regular expression: ({i, j}?) (\1) {k}, where i and j denote the minimum and maximum lengths of the motifs of the SSR sequence, respectively, and k is the minimum number of repeat motifs. From this SSR mining algorithm, we developed an SSR sequence analysis software (named "regexSSRw") that is capable of mining eligible SSR loci from FASTA format sequences; regexSSRw can be accessed at https://github.com/renm79/rgxSSRw. This SSR mining algorithm can aid a range of applications, from being used by programmers in the development of SSR mining software to being implemented by scholars into their SSR marker workflow.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- NGSpop: A desktop software that supports population studies by identifying sequence variations from next-generation sequencing data 96%
- Genome-wide DNA polymorphisms in four Actinidia arguta genotypes based on whole-genome re-sequencing 95%
- Sample Size Impact (SaSii): an R script for estimating optimal sample sizes in population genetics and population genomics studies 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- An in silico approach to identification, categorization and prediction of nucleic acid binding proteins 94%
- LDBlockShow: a fast and convenient tool for visualizing linkage disequilibrium and haplotype blocks based on variant call format files 94%
- Normalization of RNA-Seq Data using Adaptive Trimmed Mean with Multi-reference 93%
Similar papers in this journal
- Comparative analysis of machine learning and evolutionary optimization algorithms for precision tissue culture of Cannabis sativa: Prediction and validation of in vitro shoot growth and development based on the optimization of light and carbohydrate sources 93%
- Genome-wide identification, expression and bioinformatic analyses of GRAS transcription factor genes in rice 93%
- A Combinatorial Approach of Biparental QTL Mapping and Genome-Wide Association Analysis Identifies Candidate Genes for Phytophthora Blight Resistance in Sesame 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.