Novel estimators for family-based genome-wide association studies increase power and robustness
Guan, J.; Nehzati, S. M.; Benjamin, D. J.; Young, A. I.
Show abstract
A goal of genome-wide association studies (GWASs) is to estimate the causal effects of alleles carried by an individual on that individual ( direct genetic effects). Typical GWAS designs, however, are susceptible to confounding due to gene-environment correlation and non-random mating (population stratification and assortative mating). Family-based GWAS, in contrast, is robust to such confounding since it uses random, within-family genetic variation. When both parents are genotyped, a regression controlling for parental genotype provides the most powerful approach. However, parental genotypes are often missing. We have previously shown that imputing the genotypes of missing parent(s) can increase power for estimation of direct genetic effects over using genetic differences between siblings. We extend the imputation method, which previously only applied to samples with at least one genotyped sibling or parent, to singletons (individuals without any genotyped relatives). By including singletons, the effective sample size for estimation of direct effects can be increased by up to 50%. We apply this method to 408,254 White British individuals from the UK Biobank, obtaining an effective sample size increase of between 25% and 43% (depending upon phenotype) by including 368,629 singletons. While this approach maximizes power, it can be biased when there is strong population structure. We therefore introduce an imputation based estimator that is robust to population structure and more powerful than other robust estimators. We implement our estimators in the software package snipar using an efficient linear-mixed model (LMM) specified by a sparse genetic relatedness matrix. We examine the bias and variance of different family-based and standard GWAS estimators theoretically and in simulations with differing levels of population structure, enabling researchers to choose the appropriate approach depending on their research goals.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identifying Causal Variants by Fine Mapping Across Multiple Studies 97%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 96%
- Inferring Causal Direction Between Two Traits in the Presence of Horizontal Pleiotropy with GWAS Summary Data 96%
Similar papers in this journal
Similar papers in this journal
- Assumptions about frequency-dependent architectures of complex traits bias measures of functional enrichment 96%
- Statistics to prioritize rare variants in family-based sequencing studies with disease subtypes 96%
- Identifying causal genotype-phenotype relationships for population-sampled parent-child trios 96%
Similar papers in this journal
Similar papers in this journal
- Accurate modeling of replication rates in genome-wide association studies by accounting for winner's curse and study-specific heterogeneity 95%
- Comparing Heritability Estimators under Alternative Structures of Linkage Disequilibrium 95%
- Efficient approaches for large scale GWAS studies with genotype uncertainty 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.