Estimating population structure using epigenome-wide methylation data
Wang, Z.; Taylor, K.; Rotter, J.; Rich, S.; Zheng, Y.; Hou, L.; Guo, X.; Bressler, J.; Raffield, L. M.; Liu, Y.; Kaplan, R.; Lloyd-Jones, D.; Morrison, A.; Fornage, M.; Sofer, T.
Show abstract
IntroductionIn epigenome-wide association analysis (EWAS), unaddressed population stratification often leads to inflation. We aimed to compute methylation population scores (MPSs) that predict genetic principal components (GPCs) using a feature selection and regression approach. MethodsWe used multi-ethnic methylation data (Illumina 450K/EPIC array) from unrelated MESA (n=929), CARDIA (n=1123), JHS (n=1365), ARIC (n=2338), and HCHS/SOL (n=1475) individuals, randomly assigning 85% of participants from each cohort to a training dataset and the remaining 15% to a test dataset. First, we estimated the associations of GPCs with each available CpG methylation site using linear regression within each cohort, adjusting for age, sex, smoking status, race/ethnic background (as a proxy for background information associated with lifestyle and other environmental exposures that may impact methylation), alcohol use status, body mass index, and cell type proportions. We meta-analyzed the associations across cohorts and selected CpG sites with association FDR-adjusted q-value <0.05. We next aggregated individual-level data across the cohort-specific training datasets, and applied two-stage weighted least squares Lasso regression, with the GPCs as the outcomes and the selected CpG sites as penalized predictors, adjusting for the aforementioned covariates. The developed MPSs are the weighted sum of selected CpG sites from the Lasso. To evaluate the developed MPSs, we constructed them in the test dataset, and compared them with GPCs, and with MPSs constructed based on a previously-published paper. Comparison was based on correlation analysis and data visualization. We demonstrate the use of the MPSs in EWAS. ResultsIn the test dataset, the MPSs were highly correlated with GPCs, with correlation decreasing, though not monotonically, for later components. Specifically, MPS1 and GPC1 had R2= 0.99, while MPS7 and GPC7 had R2=0.27 (the lowest observed correlation). In data visualization, MPSs had similar patterns as GPCs in differentiating self-reported White, Black, and Hispanic/Latino groups, while outperforming MPC constructed using alternative published methods. MPSs showed comparable performance to GPCs in reducing some of the inflation in EWAS. ConclusionsMethylation-based population scores provide a reliable estimate of population structure in the data and can complement GPCs when genetic data are absent. Unlike previous methods based on unsupervised methylation PCA, MPSs uses supervised learning with covariate adjustment to capture genetic structure across diverse populations. The weights for each GPCs derived in our study can be applied to generate MPSs in other studies.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Elucidating the genetic architecture of DNA methylation to identify promising molecular mechanisms of disease 96%
- Expression quantitative trait methylation analysis elucidates gene regulatory effects of DNA methylation: The Framingham Heart Study 95%
- Mendelian randomization study of spermine oxidase and cancer risk 94%
Similar papers in this journal
- Pruning and thresholding approach for methylation risk scores in multi-ancestry populations 96%
- Updates to data versions and analytic methods influence the reproducibility of results from epigenome-wide association studies 94%
- CUE: CpG impUtation Ensemble for DNA Methylation Levels Across the Human Methylation450 (HM450) and EPIC (HM850) BeadChip Platforms 93%
Similar papers in this journal
- Multivariate adaptive shrinkage improves cross-population transcriptome prediction for transcriptome-wide association studies in underrepresented populations 95%
- A reference panel for linkage disequilibrium and genotype imputation using whole-genome sequencing data from 2,680 participants across India 94%
- Pathway-specific polygenic scores substantially increase the discovery of gene-adiposity interactions impacting liver biomarkers 94%
Similar papers in this journal
- A survey of rare epigenetic variation in 23,116 human genomes identifies disease-relevant epivariations and novel CGG expansions 95%
- Widespread recessive effects on common diseases in a cohort of 44,000 British Pakistanis and Bangladeshis with high autozygosity 94%
- The Phenotype-Genotype Reference Map: Improving biobank data science through replication. 94%
Similar papers in this journal
- FinaleMe: Predicting DNA methylation by the fragmentation patterns of plasma cell-free DNA 94%
- Epigenome-wide association meta-analysis of DNA methylation with coffee and tea consumption 94%
- Enhanced cell deconvolution of peripheral blood using DNA methylation for high-resolution immune profiling 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.