Back

Formal Statistical Replication Analysis in Lung Cancer Genome-Wide Association Studies

Chang, Y.-H.; Byun, J.; Gorman, B. R.; Hung, R. J.; McKay, J. D.; Amos, C. I.; Pyarajan, S.; Bhattacharya, A.; Sun, R.

2025-10-03 genetic and genomic medicine
10.1101/2025.10.02.25337130 medRxiv
Show abstract

Dozens of genome-wide association studies (GWAS) have identified thousands of single nucleotide polymor-phisms (SNPs) associated with lung cancer risk. However, it remains challenging to translate these findings to clinical insights. One well-known obstacle is the large amount of type I error attached to GWAS; attempted solutions such as setting a p-value threshold across multiple cohorts or looking for small meta-analysis p-values have only somewhat reduced false positive findings. In contrast, here we advocate for a statistical model-based replication analysis. We first demonstrate that a formal statistical test for the replication com-posite null hypothesis - i.e. that the regression coefficient of a SNP falls in the same direction in multiple cohorts simultaneously - can curate a smaller, higher-quality list of significant SNPs than common alterna-tives. In two-way simulations, the false discovery rate (FDR) of model-based replication analysis is 6.4 times lower than that of meta-analysis with a p < 10-8 threshold. In three-way replication analysis, 9.8% of the International Lung Cancer Consortium GWAS significant SNPs are replicated for squamous cell lung cancer while 33.8% are replicated for lung adenocarcinoma. Finally, we construct polygenic risk scores (PRSs) and find the replication-based PRS achieves virtually identical performance to a GWAS-significant PRS while us-ing 87.3% fewer variants. Thus, formal model-based replication analysis can greatly reduce spurious findings while still identifying important variants, allowing for more robust and more efficient translation of GWAS results.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.