FORGEPHAST: A fast, haplotype-first complex population and phenotype simulator for biobank-scale simulation
Gupta, H. V.; Raj, S. M.
Show abstract
MotivationSimulation frameworks for statistical and population genetics often separate reference-panel genotype generation, admixture modeling, phenotype simulation, and scalable genomic storage. This limits construction and evaluation of large synthetic cohorts with controlled donor composition, local ancestry, and phenotype architecture. ResultsWe developed FORGEPHAST, a Python framework for reference-panel-based synthetic haplotype and phenotype generation. FORGEPHAST supports homogeneous mosaic and pulse-admixture generation from user-defined donor panels, integrated genotype quality control, and phenotype simulation with population- and local-ancestry-specific genetic effects. Its HapStore backend provides haplotype-native, Zarr-based storage for scalable numerical analysis. We demonstrate preservation of key donor-panel population-genetic properties and use FORGEPHAST to compare polygenic risk score methods across 54 simulated phenotypic architectures at biobank scales, showing that relative performance depends on phenotype architecture, ancestry, and implementation-specific variant coverage. Availability and implementationFORGEPHAST is implemented in Python and is available as an open-source package at https://github.com/sriraj-lab/forgephast.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Efficient and Flexible Integration of Variant Characteristics in Rare Variant Association Studies Using Integrated Nested Laplace Approximation 94%
- Demonstrating the utility of flexible sequence queries against indexed short reads with FlexTyper 93%
- VUStruct: a compute pipeline for high throughput and personalized structural biology 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.