HAPNEST: efficient, large-scale generation and evaluation of synthetic datasets for genotypes and phenotypes
Wharrie, S.; Yang, Z.; Raj, V.; Monti, R.; Gupta, R.; Wang, Y.; Martin, A.; O'Connor, L. J.; Kaski, S.; Marttinen, P.; Palamara, P. F.; Lippert, C.; Ganna, A.; INTERVENE Consortium,
Show abstract
Existing methods for simulating synthetic genotype and phenotype datasets have limited scalability, constraining their usability for large-scale analyses. Moreover, a systematic approach for evaluating synthetic data quality and a benchmark synthetic dataset for developing and evaluating methods for polygenic risk scores are lacking. We present HAPNEST, a novel approach for efficiently generating diverse individual-level genotypic and phenotypic data. In comparison to alternative methods, HAPNEST shows faster computational speed and a lower degree of relatedness with reference panels, while generating datasets that preserve key statistical properties of real data. These desirable synthetic data properties enabled us to generate 6.8 million common variants and nine phenotypes with varying degrees of heritability and polygenicity across 1 million individuals. We demonstrate how HAPNEST can facilitate biobank-scale analyses through the comparison of seven methods to generate polygenic risk scoring across multiple ancestry groups and different genetic architectures.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Efficient test for deviation from Hardy Weinberg Equilibrium with known or ambiguous typing in highly polymorphic loci 96%
- kTWAS: integrating kernel-machine with transcriptome-wide association studies improves statistical power and reveals novel genes 95%
- A novel haplotype-based eQTL approach identifies genetic associations not detected through conventional SNP-based methods 95%
Similar papers in this journal
- Improving polygenic prediction from summary data by learning patterns of effect sharing across multiple phenotypes. 96%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 96%
- Genotype imputation using the Positional Burrows Wheeler Transform 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.