Removing array-specific batch effects in GWAS mega-analyses by applying a two-step imputation workflow reveals new associations for thyroid volume and goiter
Nasr, M. K.; Koenig, E.; Fuchsberger, C.; Ghasemi, S.; Voelker, U.; Voelzke, H.; Grabe, H. J.; Teumer, A.
Show abstract
BackgroundCombining individual-level data in genetic association studies (mega-analyses) enhances statistical power for identifying gene-trait associations. However, batch effects from combining variants of different arrays pose a major limitation. Here, we developed a two-step imputation workflow to overcome the array type bias. MethodsGenotype data of 10,647 individuals generated using five different arrays were included. Intermediate array-specific panels were generated and subsequently imputed against the 1000 Genomes Project Phase3 reference panel. Genetic principal component (PC) analysis assessed batch effects in the cohort-combined imputed data. The workflows performance was evaluated by comparing imputation quality r2 and allele frequency difference of the proposed two-step imputation to the conventional array-specific imputation as well as its matching with a whole-genome sequenced subgroup for further validation. We performed a genome-wide association study (GWAS) to test for genetic associations with goiter risk and thyroid gland volume, comparing summary statistics of both approaches. ResultsThe proposed workflow eliminated the batch effect from the first twenty genetic PCs. The outcome of the workflow also showed high correlation with the conventional approach for allele frequencies (r2 > 0.99). GWAS results from the two-step imputation confirmed known associations on thyroid traits and revealed novel loci for thyroid volume (TG, PAX8, IGFBP5, NRG1), and one novel locus for goiter (XKR6), which was not statistically significant following the GWAS meta-analysis of conventional imputation. ConclusionOur imputation workflow provides high-quality imputation results without technical batch effects, fostering mega-analysis involving multiple genotyping arrays for different genetic association analysis.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Imputation Disparities Driven by Recent Selectionand Their Impact on Disease Risk Estimation in East and Southeast Asian Populations 94%
- Accuracy of haplotype estimation and whole genome imputation affects complex trait analyses in complex biobanks 92%
- Direct inference and control of genetic population structure from RNA sequencing data 91%
Similar papers in this journal
- Causal effects of maternal circulating amino acids on offspring birthweight: a Mendelian randomisation study 90%
- DAGM: a novel modelling framework to assess the risk of HER2-negative breast cancer based on germline rare coding mutations 90%
- Multi-ancestry omic Mendelian randomization revealing putative drug targets of COVID-19 severity 90%
Similar papers in this journal
- The Interpretable Multimodal Machine Learning (IMML) framework reveals pathological signatures of distal sensorimotor polyneuropathy 91%
- Genome-wide DNA methylation, imprinting, and gene expression in human placentas derived from Assisted Reproductive Technology 91%
- Deep plasma proteomics identifies and validates an eight-protein biomarker panel that separate benign from malignant tumors in ovarian cancer 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.