Back

Genome-wide mining and comparative analysis of microsatellite markers from Orientia tsutsugamushi genomes

Panda, S.; Swain, S. K.; Sahu, B. P.; Sarangi, R.

2023-02-06 genomics
10.1101/2023.02.06.527248 bioRxiv
Show abstract

Microsatellite markers, otherwise known as the simple sequence repeats (SSRs), are being used for molecular identification and characterization as well as estimation of evolution pattern of the organism due to their high polymorphic nature. These are tandemly repeated sequences observed almost all organisms and differentially distributed across the genome. Although the primary genome information of Orientia tsutsugamushi (OT) suggested the repeats hold the 40% entire of its genome, but lack of characteristic of this repeats increase our interest to study more about it. Thus we investigated a genome-wide presence of microsatellites within nine complete genomes within OT and analyzed their distribution pattern, composition and complexity. The in-silico study revealed the genome of OT enrich with microsatellites having a total of 126187 SSR and 10374 cSSR throughout the genome from which 70% and 30% represented within the coding and non coding region respectively. The relative density (RD) and relative abundance (RA) of SSRs were 42-44.43/kb and 6.25-6.59/kb while for cSSRs this value ranged from 7.06-8.1/kb and 0.50-0.55/kb respectively. However, RA and RD were weakly correlate with genome size and incidence microsatellites. The mononucleotide repeats (54.55%) were prevalent over di- (33.22%), tri- (11.88%), tetra- (0.27%), penta- (0.02%), hexanucleotide (0.04%) repeats, with poly (A/T) richness over poly (G/C). Motif composition of cSSRs revealed that maximum cSSRs were made up of two microsatellites having unique duplication pattern such as AT-x-AT, CG-x-CG. More numbers microsatellites represented within the coding region provides an insight into the genome plasticity that may interfere for gene regulation to mitigate with host-pathogen interaction and evolution of the species.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Genomics
64 papers in training set
Top 0.1%
18.5%
2
PLOS ONE
5266 papers in training set
Top 14%
12.9%
3
BMC Genomics
406 papers in training set
Top 0.4%
9.7%
4
Gene
46 papers in training set
Top 0.1%
6.2%
5
Frontiers in Genetics
230 papers in training set
Top 0.4%
5.5%
50% of probability mass above
6
PeerJ
308 papers in training set
Top 1%
4.8%
7
Frontiers in Microbiology
427 papers in training set
Top 3%
4.0%
8
Gene Reports
14 papers in training set
Top 0.1%
3.2%
9
Scientific Reports
3612 papers in training set
Top 34%
3.2%
10
Genes
144 papers in training set
Top 0.9%
3.2%
11
International Journal of Biological Macromolecules
76 papers in training set
Top 0.5%
2.8%
12
Molecular Genetics and Genomics
12 papers in training set
Top 0.1%
2.1%
13
Molecular Biology Reports
21 papers in training set
Top 0.3%
2.1%
14
Current Microbiology
18 papers in training set
Top 0.6%
1.1%
15
Infection, Genetics and Evolution
42 papers in training set
Top 0.8%
0.9%
16
G3 Genes|Genomes|Genetics
351 papers in training set
Top 4%
0.8%
17
Journal of Molecular Evolution
22 papers in training set
Top 0.4%
0.8%
18
Heliyon
152 papers in training set
Top 8%
0.8%
19
Genome Biology and Evolution
338 papers in training set
Top 3%
0.8%
20
Virus Research
37 papers in training set
Top 0.8%
0.8%
21
F1000Research
88 papers in training set
Top 4%
0.8%
22
Frontiers in Plant Science
256 papers in training set
Top 4%
0.6%
23
Gigabyte
62 papers in training set
Top 1%
0.6%