Back

GenoSim: A Forward-Time Genotype Simulator for Clinical and Population Genetics with Population Stratification

Bakar, A.; Gul, R.; Haq, W. u.; Afghani, T.

2026-06-25 bioinformatics
10.64898/2026.06.20.733503 bioRxiv
Show abstract

Motivation: Next-generation sequencing studies in clinical genetics are often limited by the scarcity of human genotype data, which stems from ethical, regulatory, and economic barriers. The shortfall is sharpest in consanguineous populations, which are common in South Asia and the Middle East, where family-based designs need large pedigrees that are rarely sequenced in full. Existing simulators do not combine pedigree-aware propagation, realistic population stratification, and clinical export formats in one tool. Results: We present GenoSim, an R package for forward-time simulation of diploid SNP genotypes. It runs in two modes: a population mode implementing inbreeding-adjusted Hardy-Weinberg sampling, Wright-Fisher drift, directional selection, recurrent mutation, and Haldane recombination across multiple generations; and a pedigree-constrained mode that ingests real family VCFs and a pedigree, reconstructs phase where the pedigree makes it identifiable, propagates genotypes through the observed family structure, and appends synthetic generations. Version 1.1.1 adds population stratification through the Balding-Nichols model parameterised by gnomAD v3.1 fixation indices (F_ST) for eight ancestry groups (AFR, AMR, EAS, EUR, FIN, MID, SAS, ASJ), empirical allele-frequency loading from external reference panels, and admixed-cohort simulation. Analysis functions cover Hardy-Weinberg testing, linkage disequilibrium, runs of homozygosity, principal component analysis, founder-referenced and between-generation F-statistics, and Nei gene diversity. Availability and implementation: GenoSim is available as an R package at https://github.com/malikbak/GenoSim under the MIT licence. It requires R [≥] 4.0.0 and depends only on base R packages (stats, utils, graphics, grDevices, tools).

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

1
Bioinformatics
1204 papers in training set
Top 0.7%
26.5%
2
The American Journal of Human Genetics
234 papers in training set
Top 0.5%
9.8%
3
BMC Bioinformatics
457 papers in training set
Top 2%
5.5%
4
Genome Medicine
183 papers in training set
Top 0.6%
5.5%
5
European Journal of Human Genetics
58 papers in training set
Top 0.2%
4.8%
50% of probability mass above
6
Bioinformatics Advances
203 papers in training set
Top 1%
4.3%
7
Nature Communications
5641 papers in training set
Top 32%
4.0%
8
PLOS Computational Biology
1863 papers in training set
Top 9%
4.0%
9
Nature Genetics
286 papers in training set
Top 2%
3.4%
10
Genetics in Medicine
78 papers in training set
Top 0.5%
2.8%
11
Genetic Epidemiology
55 papers in training set
Top 0.3%
2.6%
12
PLOS Genetics
862 papers in training set
Top 5%
2.4%
13
PLOS ONE
5266 papers in training set
Top 43%
2.4%
14
Nucleic Acids Research
1281 papers in training set
Top 8%
1.9%
15
Genome Biology
637 papers in training set
Top 5%
1.7%
16
Genome Research
468 papers in training set
Top 4%
1.3%
17
Genes
144 papers in training set
Top 3%
1.1%
18
Human Genetics and Genomics Advances
84 papers in training set
Top 2%
1.0%
19
GENETICS
483 papers in training set
Top 4%
1.0%
20
BMC Medical Genomics
50 papers in training set
Top 2%
0.6%
21
BMC Genomics
406 papers in training set
Top 9%
0.6%
22
Journal of Medical Genetics
29 papers in training set
Top 0.6%
0.6%
23
Molecular Biology and Evolution
542 papers in training set
Top 5%
0.6%
24
NAR Genomics and Bioinformatics
242 papers in training set
Top 5%
0.6%