Genetic Epidemiology
○ Wiley
All preprints, ranked by how well they match Genetic Epidemiology's content profile, based on 55 papers previously published here. The average preprint has a 0.03% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.
Chen, T.; Voorhies, K.; Reeson, A.; Seo, S.; Lee, S.; Hahn, G.; Hecker, J.; Prokopenko, D.; Hoth, K.; Kelly, R.; Lasky-Su, J. A.; Weiss, S.; Lange, C.; Lutz, S.
Show abstract
Mendelian Randomization (MR) is a popular tool for inferring causal relationships between traits using genetic variants as instrumental variables. These methods have been extended to also determine the direction of causality. However, causal direction cannot be inferred from a statistical test or estimation procedure (i.e. from data alone) without further assumptions and the methods operating characteristics and relative performances are not well understood. We conducted a comprehensive simulation study to illustrate this issue by evaluating type I error and power of 17 summary-based MR methods for inferring the effect direction. These methods fall within three methodological families: MR Steiger, Causal Direction (CD), and bidirectional MR approaches, with scenarios ranging across combinations of horizontal pleiotropy, unmeasured confounding, measurement error, longitudinal feedback, and varying sample sizes. While most methods achieved sufficient power levels under the alternative hypothesis in most scenarios, we found that every method was susceptible to inferring the wrong causal direction or under powered, and no method consistently maintained both correct type 1 error control and high power. In our applications, we evaluated the effect direction between the trait pairs body mass index (BMI) and major depressive disorder (MDD) and between BMI and asthma. To help researchers to evaluate the 17 methods to infer the effect direction and consider these challenges in their own data, we have developed MRdirection, an R package that runs the simulation studies examining the 17 directional MR methods across different user-defined scenarios. Our study, together with the accompanying R package, provides researchers with a tool for examining directional MR methods given different underlying assumptions.
Mukhopadhyay, N.; Feingold, E. E.; Brand, H.; Lee, M. K.; Kurtas, E. N.; Sanchis-Juan, A.; Moreno-Uribe, L.; Wehby, G.; Valencia-Ramirez, L. C.; Restrepo Muneton, C. P.; Padilla, C.; Deleyiannis, F.; Poletta, F. A.; Orioli, I. M.; Hecht, J. T.; Buxo, C. J.; Butali, A.; Adeyemo, W. L.; Abebe, M. E.; Vieira, A. R.; Shaffer, J. R.; Murray, J. C.; Weinberg, S. M.; Ruczinski, I.; Leslie-Clarkson, E. J.; Marazita, M. L.
Show abstract
ObjectiveOur understanding of the genetic causes of non-syndromic orofacial clefts (OFCs) is based largely upon genetic studies of common and rare nucleotide variants. Less is known about the role of copy number variations (CNVs) and the studies published to date have been limited to either small samples or targeted genomic regions. The objective of our study is to investigate the contribution of CNVs spread across the entire genome to OFC risk in a large multi-ancestry cohort. MethodsWe utilized PennCNV on microarray genotyping data to detect CNVs in 10,240 participants (2,484 with clefts, 7,756 unaffected). 70,695 quality-filtered autosomal CNVs (49,660 deletions, 21,035 duplications) were used to assign normal/abnormal copy number statuses at 67,199 positions from the GRCh37 genome assembly. Genome-wide association was run between cleft status and copy number status. ResultsWe observed a highly significant association between OFCs and deletions on chromosome 7p14.1 (p=1.32e-35) driven by Central and South American ancestry (p=1.04e-25) participants, with less significant contributions from European (p=3.37e-08) and Asian (p=0.01) ancestry participants. We also observed four other loci with p-values below 10e-04. ConclusionThe 7p14.1 association observed in our study is a replication of two prior studies in independent cohorts of European ancestry. However, this locus lies in a T-cell receptor region that is subject to somatic rearrangements that decrease in frequency with age and may affect genetic association results. Our data show age effects as well as differences between blood and saliva samples. Thus, our results can be interpreted either as supporting a previously established association with orofacial clefts, or as questioning those previous results in favor of a hypothesis about the behavior of somatic rearrangements in T-cell receptor regions.
Lewinger, J. P.; Kawaguchi, E. S.; Gauderman, W. J.
Show abstract
Two-step testing has emerged as the leading approach for genome-wide gene-environment interaction scans, providing greater power than standard single-step methods in nearly all plausible scenarios. This gain in power has enabled several recent discoveries in gene-environment interactions. However, despite controlling the genome-wide type I error rate, two-step tests typically lack valid p-values, hindering comparisons with single-step results. We demonstrate how multiple-testing adjusted p-values can be defined for two-step tests using standard multiple-testing theory, and how these can be further scaled to allow valid comparisons with single-step tests.
Li, Y.; Lei, H.; Wen, X.; Cao, H.
Show abstract
Replicability is the cornerstone of modern scientific research. Reliable identifications of genotype-phenotype associations that are significant in multiple genome-wide association studies (GWASs) provide stronger evidence for the findings. Current replicability analysis relies on the independence assumption among single nucleotide polymorphisms (SNPs) and ignores the linkage disequilibrium (LD) structure. We show that such a strategy may produce either overly liberal or overly conservative results in practice. We develop an efficient method, ReAD, to detect replicable SNPs associated with the phenotype from two GWASs accounting for the LD structure. The local dependence structure of SNPs across two heterogeneous studies is captured by a four-state hidden Markov model (HMM) built on two sequences of p-values. By incorporating information from adjacent locations via the HMM, our approach provides more accurate SNP significance rankings. ReAD is scalable, platform independent and more powerful than existing replicability analysis methods with effective false discovery rate (FDR) control. Through analysis of datasets from two asthma GWASs and two ulcerative colitis GWASs, we show that ReAD can identify replicable genetic loci that existing methods might otherwise miss.
Wolf, J. M.; Westra, J.; Tintle, N.
Show abstract
While the promise of electronic medical record and biobank data is large, major questions remain about patient privacy, computational hurdles, and data access. One promising area of recent development is pre-computing non-individually identifiable summary statistics to be made publicly available for exploration and downstream analysis. In this manuscript we demonstrate how to utilize pre-computed linear association statistics between individual genetic variants and phenotypes to infer genetic relationships between products of phenotypes (e.g., ratios; logical combinations of binary phenotypes using and and or) with customized covariate choices. We propose a method to approximate covariate adjusted linear models for products and logical combinations of phenotypes using only pre-computed summary statistics. We evaluate our methods accuracy through several simulation studies and an application modeling various fatty acid ratios using data from the Framingham Heart Study. These studies show consistent ability to recapitulate analysis results performed on individual level data including maintenance of the Type I error rate, power, and effect size estimates. An implementation of this proposed method is available in the publicly available R package pcsstools.
de la Fuente, J.; Londono-Correa, D.; Tucker-Drob, E. M.
Show abstract
Within the Genomic SEM framework, common factors are often used to index shared genetic etiology across constellations of GWAS phenotypes. A standard common pathway model, in which a genetic association is estimated between an external GWAS phenotype and a common factor, assumes that all genetic associations between the external GWAS phenotype and the individual indicator phenotypes are mediated through the factor. This assumption can be tested using the QTrait statistic, which compares the common pathway model to an independent pathways model that allows for direct genetic associations between the external GWAS phenotype and the individual indicators of the factor. We expand upon the QTrait approach by describing an effect size index that quantifies the degree to which the common pathways model is violated, and we provide a systematic approach for empirically identifying specific direct pathways between an external trait and indicator traits. Our method comprises a series of omnibus tests and outlier detection algorithms indexing the heterogeneity of associations between the genetic component of external traits and the individual indicators of common factors. We provide a set of automated functions which we apply to investigate the patterns of genetic associations across a set of external correlates with respect to indicators of general cognitive ability and case-control and proxy GWAS indices of Alzheimers disease.
Gauderman, W. J.; Fu, Y.; Quem, B.; Kawaguchi, E.; Wang, Y.; Morrison, J.; Brenner, H.; Chan, A.; Gruber, S.; Temitope, K.; Li, L.; Moreno, V.; Pellatt, A.; Peters, U.; Samadder, N. J.; Schmit, S.; Ulrich, C.; Um, C.; Wu, A.; Lewinger, J. P.; Mi, H.; Drew, D.
Show abstract
A polygenic risk score (PRS) is used to quantify the combined disease risk of many genetic variants. For complex human traits there is interest in determining whether the PRS modifies, i.e. interacts with, important environmental (E) risk factors. Detection of a PRS by environment (PRS x E) interaction may provide clues to underlying biology and can be useful in developing targeted prevention strategies for modifiable risk factors. The standard PRS may include a subset of variants that interact with E but a much larger subset of variants that affect disease without regard to E. This latter subset will water down the underlying signal in former subset, leading to reduced power to detect PRS x E interaction. We explore the use of pathway-defined PRS (pPRS) scores, using state of the art tools to annotate subsets of variants to genomic pathways. We demonstrate via simulation that testing targeted pPRS x E interaction can yield substantially greater power than testing overall PRS x E interaction. We also analyze a large study (N=78,253) of colorectal cancer (CRC) where E = non-steroidal anti-inflammatory drugs (NSAIDs), a well-established protective exposure. While no evidence of overall PRS x NSAIDs interaction (p=0.41) is observed, a significant pPRS x NSAIDs interaction (p=0.0003) is identified based on SNPs within the TGF-{beta} / gonadotropin releasing hormone receptor (GRHR) pathway. NSAIDS is protective (OR=0.84) for those at the 5th percentile of the TGF-{beta}/GRHR pPRS (low genetic risk, OR), but significantly more protective (OR=0.70) for those at the 95th percentile (high genetic risk). From a biological perspective, this suggests that NSAIDs may act to reduce CRC risk specifically through genes in these pathways. From a population health perspective, our result suggests that focusing on genes within these pathways may be effective at identifying those for whom NSAIDs-based CRC-prevention efforts may be most effective. Author SummaryThe identification of polygenic risk score (PRS) by environment (PRSxE) interactions may provide clues to underlying biology and facilitate targeted disease prevention strategies. The standard approach to computing a PRS likely includes many variants that affect disease without regard to E, reducing power to detect PRS x E interactions. We utilize gene annotation tools to develop pathway-based PRS (pPRS) scores and show by simulation studies that testing pPRS x E interaction can yield substantially greater power than testing PRS x E, while also integrating biological knowledge into the analysis. We apply our method to a large study of colorectal cancer to identify a significant pPRS x NSAIDs interaction (p=0.0003) based on SNPs within the TGF-{beta} / gonadotropin releasing hormone receptor (GRHR) pathway. Our findings suggest that focusing on genetic susceptibility within biologically informed pathways may be more sensitive for identifying exposures that can be considered as part of a precision prevention approach.
Zhang, M.; Tang, J.; Brown, M. R.; Morrison, A. C.; Boerwinkle, E.; Manning, A. K.; Liu, C.-T.; Chen, H.
Show abstract
Linear mixed models (LMMs) are widely used in gene-environment interaction (GEI) studies to account for population structure and relatedness. However, genome-wide GEI tests using LMMs are computationally intensive, and model-based tests can yield inflated type I error rates when environmental main effects are misspecified. While robust inference methods exist for unrelated samples, challenges remain for related individuals. A common workaround is a two-step approach that first adjusts for relatedness via an LMM and then uses residuals in a standard linear model, but its validity for GEI studies is unclear. We propose a robust mixed model association test (RoM) for large-scale GEI analysis in related samples. RoM uses the Huber-White sandwich estimator and offers efficient computation, scaling linearly with sample size when cluster sizes are bounded. Simulations show that RoM achieves better type I error control at genome-wide significance levels than both the two-step method and alternative strategies. We apply RoM to GEI analyses of waist-hip ratio (WHR) with BMI using data from the Framingham Heart Study (7,264 related individuals), ARIC (9,312 individuals with repeated measures), and WHR with sex using data from UK Biobank (407,068 related individuals), confirming robust error control and comparable signal detection.
Lin, B.; Sun, L.
Show abstract
The effect of a genetic variant on a complex trait may differ between female and male, and in the presence of such genetic effect heterogeneity, sex-stratified analysis is often used. For example, genetic effects are sex-specific for testosterone levels, and sex-stratified analysis of testosterone in literature provided easy-to-interpret, sex-specific effect size estimates. However, from the perspective of association testing power, sex-stratified analysis may not be the best approach. As sex-specific genetic effect implies SNPxSex interaction effect, jointly testing SNP main and SNPxSex interaction effects may be more powerful than sex-stratified analysis or the standard main-effect testing approach. Moreover, since individual data may be unavailable, it is then of interest to study if the interaction analysis can be derived from sex-stratified summary statistics. We considered several different sex-combined methods and evaluated them through extensive simulation studies. We observed that a) the joint SNP main and SNPxSex interaction analysis is most robust to a wide range of genetic models, and b) this joint interaction testing result can be obtained by quadratically combining sex-stratified summary statistics (i.e. squared sum of the sex-stratified summary statistics). We then reanalysed the testosterone levels of the UK Biobank data using sex-combined interaction analysis, which identified 27 new loci that were missed by the sex-stratified approach and the standard sex-combined analysis. Finally, we provide supporting association evidence for nine new loci, uniquely identified by the sex-combined interaction analysis, from earlier association studies of either testosterone level or steroid biosynthesis pathway where testosterone is synthesized. We thus recommend sex-combined interaction analysis, particularly for traits with known sex differences, for most powerful association testing, then followed by sex-stratified analysis for effect size estimation and interpretation.
Patel, A.; Gill, D.; Shungin, D.; Mantzoros, C. S.; Knudsen, L. B.; Bowden, J.; Burgess, S.
Show abstract
Phenotypic heterogeneity at genomic loci encoding drug targets can be exploited by multivariable Mendelian randomization to provide insight on the pathways by which pharmacological interventions may affect disease risk. However, statistical inference in such investigations may be poor if overdispersion heterogeneity in measured genetic associations is unaccounted for. In this work, we first develop conditional F-statistics for dimension-reduced genetic associations that enable more accurate measurement of phenotypic heterogeneity. We then develop a novel extension for two-sample multivariable Mendelian randomization that accounts for overdispersion heterogeneity in dimension-reduced genetic associations. Our empirical focus is to use genetic variants in the GLP1R gene region to understand the mechanism by which GLP1R agonism affects coronary artery disease (CAD) risk. Colocalization analyses indicate that distinct variants in the GLP1R gene region are associated with body mass index and type 2 diabetes. Multivariable Mendelian randomization analyses that were corrected for overdispersion heterogeneity suggest that bodyweight lowering rather than type 2 diabetes liability lowering effects of GLP1R agonism are more likely contributing to reduced CAD risk. Tissue-specific analyses prioritised brain tissue as the most likely to be relevant for CAD risk, of the tissues considered. We hope the multivariable Mendelian randomization approach illustrated here is widely applicable to better understand mechanisms linking drug targets to diseases outcomes, and hence to guide drug development efforts.
Gao, J.; Sun, L.
Show abstract
Functional annotations have the potential to increase the power of genome-wide association studies (GWAS) by prioritizing variants according to their biological function. Focusing on variant-specific annotation meta-scores including CADD (Kircher et al., 2014) and Eigen (Ionita-laza et al., 2016), we broadly examined GWAS summary statistics of 1,132 traits from the UK Biobank (Sudlow et al., 2015) using the weighted p-value approach (Genovese et al., 2006) and stratified false discovery control (sFDR) method (Sun et al., 2006). These 1,132 traits were rated by Benjamin Neales lab from the Broad Institute as having medium to high confidence for their heritability estimates. Averaged across the 1,132 UK Biobank traits, sFDR was more robust to uninformative meta-scores, but the weighted p-value method identified more variants using CADD or Eigen, based on performance measures that included type I error control, recall, precision, and relative efficiency. Our application results were consistent with those from an extensive simulation study using three different designs, including leveraging the real genetic data combined with simulated genomic data and vice versa. We also considered the recent FINDOR method (Kichaev et al., 2019), which leverages a set of individual 75 functional annotations into GWAS. An earlier application of FINDOR to 27 traits selected from the z7 category (SNP-heritability p-value < 1.27 x 10-12 by Nealelab) detected 13%-20% additional genome-wide significant loci as compared to the standard annotation-free GWAS, which we confirmed. Moreover, across all 438 traits in the z7 category, 46,631 out 59,764 (80%) significant loci discovered are common across the three data-integration methods. However, across all the 1,132 UK Biobank traits examined, the median [Q1,Q3] of the total numbers of new, genome-wide significant independent loci were 0 [0, 3] by FINDOR, 0 [0, 2] by weighted p-value, and 0 [0, 0] by sFDR. Notably, 162 traits (89%) in the nonsig trait category (SNP-heritability p-value > 0.05, "likely reflecting limited statistical power rather than a true lack of heritability" by Nealelab) had no new discoveries after data-integration by any of the three methods. Our findings suggest that more informative scores or new data integration methods are warranted to further improve the power of GWAS by leveraging the variant functional annotations.
Shpak, M.; Parfitt, E.; Mahmoudiandehkordi, S.; Maadooliat, M.; Schrodi, S. J.
Show abstract
Common diseases exhibit substantial heritability, and GWAS of these diseases have revealed hundreds of thousands of high-frequency disease susceptibility variants throughout the genome. These studies offer the prospect of using genomic data to improve disease prediction and diagnosis, however, the relative performance of different predictive modeling approaches is not well-characterized. To investigate this systematically, we constructed a Monte Carlo simulation generating model genomes with large numbers of SNPs, with a proportion of SNPs carrying risk alleles that are parameterized by the strength of their effects and by different modes of inheritance - additive, dominant, recessive, and combinations thereof. After generating genotypes for cases and controls, several machine learning classifiers (logistic regression, naive Bayes, random forests, and neural networks, with and without feature selection) were applied to predict disease phenotype from genotypes. Each classifiers rates of false positives and false negatives were evaluated and compared using AUC. We found that random forest models were the most accurate predictors of disease phenotype over the range of inheritance parameters, followed by logistic regression and naive Bayes, while the feedforward multilayer neural network-based predictive model had lower AUC. Furthermore, with the small fraction of null sites in our model, there was almost no difference in the performance of classifiers with or without LASSO-based feature selection. We also investigate the association of AUC with the difference in polygenic risk score (PRS) between disease and control samples by comparing AUC in the simulations to the values predicted from the PRS distributions based on odds-risk and liability models.
Tyrer, J.; Peng, P.-C.; DeVries, A.; Gayther, S.; Jones, M. R.; Pharoah, P.
Show abstract
Structured AbstractO_ST_ABSMotivationC_ST_ABSAs precision medicine advances, polygenic scores (PGS) have become increasingly important for clinical risk assessment. Many methods have been developed to create polygenic models with increased accuracy for risk prediction. Our select and shrink with summary statistics (S4) PGS method extends a previous method (polygenic risk score - continuous shrinkage (PRS-CS)) by using a continuous shrinkage prior on effect sizes with a selection strategy for including SNPs to create the best performing model. ResultsThe S4 method provides overall improved PGS accuracy for UK Biobank participants when compared to LDpred2 and PRS-CS across a variety of phenotypes with differing genetic architectures. Additionally, the S4 method has higher estimated PGS accuracy over LDpred2 in Finnish and Japanese populations. Thus, the S4 method represents an improvement in overall PGS accuracy across multiple phenotypes and increases the transferability of PGS across ancestries. Availability and ImplementationThe S4 program is freely available at https://github.com/jpt34/S4_programs. Supplementary informationSupplementary data [will be] available at Bioinformatics online.
Romanescu, R.; Liu, M.
Show abstract
We consider the problem of optimal testing for genetic interaction between two variants, allowing for possible main effects. Finding a most powerful test is important because it ends a series of attempts in the literature to construct ever more powerful tests for interaction at the variant pair level. Testing under a logistic regression model is known to be underpowered, partly because patterns of enrichment in the genotypes themselves are lost when regarding genotypes solely as predictors. Instead, we use the retrospective likelihood approach, which makes use of all the data by treating genotypes as outcomes alongside affection status. Using a parsimonious parameterization of penetrance based on the risk ratio, which links directly to the population prevalence and avoids having to estimate an intercept term, we construct an approximate uniformly most powerful unbiased test for interaction. This test is based on optimal testing theory and accounts for nuisance main effects without requiring their explicit estimation. The test statistic can be easily modified for optimal testing under other modes of genetic interaction, such as recessive x recessive or dominant x dominant. We demonstrate significant power gains compared to the odds-ratio-based PLINK test, in simulation studies. Finally, we apply the test to scan for interactions in IBD cases and controls from the UK Biobank. The top SNP pairs show enrichment for a pathway related to existing therapies for IBD.
Dugan, A. J.; Fardo, D. W.; Zaykin, D. V.; Vsevolozhskaya, O. A.
Show abstract
Genetic pleiotropy is the phenomenon where a single gene or genetic variant influences multiple traits. Numerous statistical methods exist for testing for genetic pleiotropy at the variant level, but fewer methods are available for testing genetic pleiotropy at the gene-level. In the current study, we derive an exact alternative to the Shen and Faraway functional F-statistic for functional-on-scalar regression models. Through extensive simulation studies, we show that this exact alternative performs similarly to the Shen and Faraway F-statistic in gene-based, multi-phenotype analyses and both F-statistics perform better than existing methods in small sample, modest effect size situations. We then apply all methods to real-world, neurodegenerative disease data and identify novel associations.
Dong, R.; Wang, M.; Wang, G. T.; deWan, A. T.; Leal, S. M.
Show abstract
Motivation: Linkage disequilibrium score (LDSC) regression is a popular method to estimate heritability for complex traits using summary statistics and linkage disequilibrium (LD) reference panels, offering a practical alternative to methods requiring individual-level data. Despite its widespread use, LDSC regression can produce biased heritability estimates. The properties of LDSC regression were investigated using summary statistics from several large-scale Alzheimer's disease (AD) studies and a variety of LD reference panels. These heritability estimates were compared with those obtained from individual-level data. Results: When LDSC regression was applied to summary statistics obtained from meta-analysis, it led to an underestimation of heritability. This can occur if meta-analysis is used to combine studies of different ancestries leading to the caveat of the lack of an appropriate LD reference panel. Additionally meta-analyses often include studies with different phenotype definitions, that not only impacts heritability estimates but also makes them uninterpretable. Summary statistics generated from imputed variants, even those with high imputation accuracy, can lead to underestimation of heritability. For example, the heritability estimates for AD were reduced from 0.265 (se 0.148) to 0.160 (se 0.041) when imputed variants (INFO>0.9) were included compared to analyzing only genotype array variants. A decrease in heritability estimates was also observed when individual-level imputed variant data were analyzed using GCTA-GREML. Our findings highlight the caveats of estimating heritability using meta-analysis summary statistics or imputed data instead of genotyped or sequence data.
Hernandez Cordero, A. I.; Milne, S.; Yang, C. X.; Li, X.; Shi, H.; Sin, D. D.; Obeidat, M.
Show abstract
BackgroundLarge genome-wide association studies (GWAS) and other genetic studies have revealed genetic loci that are associated with chronic obstructive pulmonary disease (COPD). However, the proteins responsible for COPD pathogenesis remain elusive. We used integrative-omics by combining genetics of lung function and COPD with genetics of proteome to identify proteins underlying lung function variation and COPD risk. MethodsWe used summary statistics from the GWAS of human plasma proteome from the INTERVAL cohort (n=3,301) and integrated these data with lung function GWAS results from the UK Biobank cohorts (n=400,102) and COPD GWAS results from the ICGC cohort (35,735 cases and 222,076 controls). We performed in parallel: a proteome-wide Bayesian colocalization, and a proteome-wide Mendelian Randomization (MR) analyses. Next, we selected proteins that colocalized with lung function and/or COPD risk and explored their causal association with lung function and/or COPD using MR analysis (P<0.05). ResultsWe found 537, 607, and 250 proteins that colocalized with force expiratory volume in one second (FEV1), FEV1/forced vital capacity (FVC), or COPD risk, respectively. Of these, 1,051 were unique proteins. The sRAGE protein demonstrated the strongest colocalization with FEV1/FVC and COPD risk, while QSOX2, FAM3D and F177A proteins had the strongest associations with FEV1. Of these, 37 proteins that colocalized with lung function and/or COPD, also had a significant causal association. These included proteins such as PDE4D, QSOX2 and RGAP1, amongst others. ConclusionIntegrative-omics reveals new proteins related to lung function. These proteins may play important roles in the pathogenesis of COPD.
Zhang, L.; Sun, L.
Show abstract
In a case-control association study, deviation from Hardy-Weinberg equilibrium (HWE) or Hardy-Weinberg dis-equilibrium (HWD) in the control group is usually considered as evidence for potential genotyping error, and the corresponding SNP is then removed from the study. On the other hand, assuming HWE holds in the study population, a truly associated SNP is expected to be out of HWE in the case group. Efforts have been made in combining association tests with tests of HWE in the cases to increase the power of detecting disease susceptibility loci (Song and Elston (2006), Wang and Shete (2010)). However, these existing methods are ad-hoc and sensitive to model assumptions. Utilizing the recent robust allele-based (RA) regression model for conducting allelic association tests (Zhang and Sun (2020)), here we propose a joint RA test that naturally integrates association evidence from the traditional association test and a test that evaluates the difference in HWD between the case and control groups. The proposed test is robust to genotyping error, as well as to potential HWD in the population attributed to factors that are unrelated to phenotype-genotype association. We provide the asymptotic distribution of the proposed test statistic so that it is easy to implement, and we demonstrate the accuracy and efficiency of the test through extensive simulation studies and an application.
Bureau, A.; Girard, S.; Moreau, C.; Maziade, M.; Oubninte, S.
Show abstract
AbstractThe missing heritability caused by rare variants (RVs) poses a significant challenge to pre-established statistical methods. Our study aims at detecting RVs using identical-by-descent (IBD) segments as a proxy for recent variants in family data from a population with a founder effect for which genealogy is available--a distinguishing feature of our approach. Inferring IBD segments from genotype array data, which is more accessible than whole genome sequences, enables application to large sample sizes. Our approach involves dividing the genome into fixed-length windows, treating each window as a synthetic genomic region (SG), and then identifying groups of affected individuals sharing a specific IBD segment over an SG by analyzing genotype array data to infer pairwise IBD segments. Data from pairwise IBD segments is then used to identify densely connected haplotypes as IBD clusters via DASH. Lastly, we adapt, implement, and evaluate statistics to test for IBD sharing enrichment among affected individuals within SGs. The null distribution of the genome-wide maximal statistic value is obtained by simulating whole-genome transmission in a genealogy using msprime. For application purposes, Eastern Quebec has been studied as an example of a population with a founder effect. Using the BALSAC database to reconstruct the genealogy of 1,200 subjects across 48 schizophrenia and bipolar disorder multi-generational families led to an 18-generation pedigree with 84% completeness at the 10th generation. The statistic denoted as Smsg for the "most shared haplotype in an SG" exhibits superior power in detecting causal SGs when compared to the adapted Sall measure and (with a single causal variant in a region) to GMMAT (Generalized Linear Mixed Model Association Test) applied to IBD clusters. Our analysis of data pertaining to schizophrenia and bipolar disorder reveals no regions that surpass the conventional significance thresholds for harboring rare variants associated with these disorders. Two distinct regions--on chromosomes 5 and 11--stand out due to their maximal Smsg values. These findings underscore the potential of leveraging genealogical data and IBD segments to uncover rare variants in complex diseases.
Moon, H.; Sloofman, L.; Avila, M. N.; Klei, L.; Devlin, B.; Buxbaum, J.; Roeder, K.
Show abstract
MotivationGene-damaging mutations are highly informative for studies seeking to discover genes underlying developmental disorders. Traditionally, these de novo variants are recognized by evaluating high-quality DNA sequence from affected offspring and parents. However, when parental sequence is unavailable, methods are required to infer de novo status and use this inference for association studies. ResultsWe use data from autism spectrum disorder to illustrate and evaluate methods. Separating de novo from rare inherited variants is challenging because the latter are far more common. Using a classifier for unbalanced data and variants of known inheritance class, we build an inheritance model and then a de novo score for variants when parental data are missing. Next, we propose a new Random Draw (RD) model to use this score for gene discovery. Built into an existing inferential framework, RD produces a more powerful gene-based association test and controls the false discovery rate. Availability and ImplementationThe implementation code and publicly available data are provided at: https://github.com/HaeunM/TADA-RD.