Formal Statistical Replication Analysis in Lung Cancer Genome-Wide Association Studies
Chang, Y.-H.; Byun, J.; Gorman, B. R.; Hung, R. J.; McKay, J. D.; Amos, C. I.; Pyarajan, S.; Bhattacharya, A.; Sun, R.
Show abstract
Dozens of genome-wide association studies (GWAS) have identified thousands of single nucleotide polymor-phisms (SNPs) associated with lung cancer risk. However, it remains challenging to translate these findings to clinical insights. One well-known obstacle is the large amount of type I error attached to GWAS; attempted solutions such as setting a p-value threshold across multiple cohorts or looking for small meta-analysis p-values have only somewhat reduced false positive findings. In contrast, here we advocate for a statistical model-based replication analysis. We first demonstrate that a formal statistical test for the replication com-posite null hypothesis - i.e. that the regression coefficient of a SNP falls in the same direction in multiple cohorts simultaneously - can curate a smaller, higher-quality list of significant SNPs than common alterna-tives. In two-way simulations, the false discovery rate (FDR) of model-based replication analysis is 6.4 times lower than that of meta-analysis with a p < 10-8 threshold. In three-way replication analysis, 9.8% of the International Lung Cancer Consortium GWAS significant SNPs are replicated for squamous cell lung cancer while 33.8% are replicated for lung adenocarcinoma. Finally, we construct polygenic risk scores (PRSs) and find the replication-based PRS achieves virtually identical performance to a GWAS-significant PRS while us-ing 87.3% fewer variants. Thus, formal model-based replication analysis can greatly reduce spurious findings while still identifying important variants, allowing for more robust and more efficient translation of GWAS results.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Leveraging expression from multiple tissues using sparse canonical correlation analysis (sCCA) and aggregate tests improves the power of transcriptome-wide association studies (TWAS) 96%
- Adjusting for principal components can induce spurious associations in genome-wide association studies in admixed populations 95%
- Beyond SNP Heritability: Polygenicity and Discoverability of Phenotypes Estimated with a Univariate Gaussian Mixture Model 95%
Similar papers in this journal
- Assumptions about frequency-dependent architectures of complex traits bias measures of functional enrichment 95%
- Using Family History Data to Improve the Power of Association Studies: Application to Cancer in UK Biobank 95%
- Identity-by-descent mapping using multi-individual IBD with genome-wide multiple testing adjustment 94%
Similar papers in this journal
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 96%
- Identifying tumor cells at the single cell level 93%
- Bayesian Multi-Study Non-Negative Matrix Factorization for Mutational Signatures 93%
Similar papers in this journal
- Accounting for genetic effect heterogeneity in fine-mapping and improving power to detect gene-environment interactions with SharePro 96%
- SUMMIT: An integrative approach for better transcriptomic data imputation improves causal gene identification 95%
- Fast Kernel-based Association Testing of non-linear genetic effects for Biobank-scale data 95%
Similar papers in this journal
- Pitfalls in performing genome-wide association studies on ratio traits 95%
- Inverted genomic regions between reference genome builds in humans impact imputation accuracy and decrease the power of association testing 94%
- Disease-specific prioritization of non-coding GWAS variants based on chromatin accessibility 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.