Widespread genetic effect heterogeneity impacts bias and power in nonlinear Mendelian randomization
Wang, J.; Morrison, J.
Show abstract
1Mendelian randomization (MR) uses genetic variants as instrumental variables to infer causal relationships between complex traits. Standard MR can be used to estimate an average causal effect at the population level, and typically assumes a linear exposure-outcome relationship. Recently, several methods for estimating nonlinear effects have been developed. However, many have been found to produce spurious empirical findings when subjected to negative control analyses. We propose that this poor performance may be attributable to heterogeneity in variant-exposure associations. We demonstrate that heterogeneous genetic effects on exposure lead to biased estimates, poor coverage, and inflated type I error in control function and stratification-based methods. In contrast, two-stage least squares (TSLS) methods are robust to such heterogeneity, but suffer from low precision and low power in some circumstances. We show that a statistical test for heterogeneity can be used to guide the choice of nonlinear MR methods. Using UK Biobank data, we reassess the causal effects of BMI, vitamin D, and alcohol consumption on blood pressure, lipid, C-reactive protein, and age (negative control). We find strong evidence of heterogeneity for all three exposures, and also recapitulate previous results that control function and stratification-based methods are prone to false positives. Finally, using nonparametric TSLS, we identify evidence of nonlinear causal effects of BMI on HDL cholesterol, triglycerides, and C-reactive protein; however, specific estimates of the shape of these relationships are imprecise. Altogether, our results suggest that common nonlinear MR methods are unreliable in the presence of realistic levels of heterogeneity, and that more methodological development is required before practically useful nonlinear MR is feasible.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking Mendelian Randomization methods for causal inference using genome-wide association study summary statistics 97%
- Welch-weighted Egger regression reduces false positives due to correlated pleiotropy in Mendelian randomization 97%
- A novel and efficient machine learning Mendelian randomization estimator applied to predict the safety and efficacy of sclerostin inhibition 96%
Similar papers in this journal
- A Comprehensive Evaluation of Methods for Mendelian Randomization Using Realistic Simulations and an Analysis of 38 Biomarkers for Risk of Type-2 Diabetes 96%
- Bias in two-sample Mendelian randomization when using heritable covariable-adjusted summary associations 96%
- Estimation of time-varying causal effects with multivariable Mendelian randomization: some cautionary notes 94%
Similar papers in this journal
- A Hierarchical Approach Using Marginal Summary Statistics for Multiple Intermediates in a Mendelian Randomization or Transcriptome Analysis 94%
- A structural mean modelling Mendelian randomization approach to investigate the lifecourse effect of adiposity: applied and methodological considerations 92%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 92%
Similar papers in this journal
- Simultaneous estimation of bi-directional causal effects and heritable confounding from GWAS summary statistics 97%
- Causal mediation analysis for time-varying heritable risk factors with Mendelian Randomization 96%
- Accounting for genetic effect heterogeneity in fine-mapping and improving power to detect gene-environment interactions with SharePro 96%
Similar papers in this journal
- Mendelian Randomization with longitudinal exposure data: simulation study and real data application 96%
- Efficient Estimation of Indirect Effects in Case-Control Studies Using a Unified Likelihood Framework 95%
- Bayesian Variable Selection with a Pleiotropic Loss Function in Mendelian Randomization 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.