RAREsim2: Flexible simulation of rare variant genetic data using real haplotypes
Murphy, J. I.; Barnard, R.; Null, M. M.; Hendricks, A. E.
Show abstract
SummaryRealistic simulated data is critical for advancing methodological development and optimizing study design in genetics research. However, many genetic simulation tools are unable to replicate the distribution of rare variants or incorporate key genetic information, such as functional annotations and linkage disequilibrium. RAREsim, an accurate rare-variant simulation algorithm that uses real genetic haplotypes, was developed to address these limitations. Here, we introduce RAREsim2, an update that provides both streamlined software and new functionalities for simulating individual-level differences (e.g., cases vs controls, technological or batch effects) and variant-level differences to represent a variety of causal models. We demonstrate RAREsim2s utility with three rare variant association methods (Burden, SKAT, and SKAT-O) across several simulation scenarios: causal variants only in cases, causal variants in both cases and controls, and no causal variants. Type I Error was maintained and the optimal test matched previously known patterns: Burden performed best given a large amount of causal variants with the same direction of effect; SKAT performed best given causal variants with opposite directions of effect. SKAT-O was powerful across all simulation scenarios. We highlight RAREsim2s capabilities to simulate various genetic ancestries (African, East Asian, Non-Finnish European, and South Asian), gene sizes (~20-80 functional rare variants per gene), strengths of association (20%, 40%, 60% functional variants), and proportions of risk variants (1, 0.75, 0.5). Importantly, real genetic regions can be simulated to include known variant functions and disease associations. Ultimately, RAREsim2 offers additional flexibility and ease in simulating a multitude of realistic genetic scenarios. Availability and implementationThe RAREsim2 python package is available at https://github.com/Hendricks-Research-Team/RAREsim2 and the code for the example demonstration is available at https://github.com/JessMurphy/RAREsim2_demo. Contactjessica.murphy@cuanschutz.edu
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- CLUES2 Companion: Computational pipelines to estimate, visualize, and date selection on multi-locus sites 95%
- HTRX: an R package for learning non-contiguous haplotypes associated with a phenotype 94%
- RegionScan: A comprehensive R package for region-level genome-wide association testing with integration and visualization of multiple-variant and single-variant hypothesis testing 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.