Pragmatic vs. naive genetic instrument selection in Mendelian randomization studies: a practical guide
Mason, A. C.; Ballabio, G.; Paz, V.; Sofat, R.; Garfield, V.
Show abstract
Mendelian randomization (MR) is widely used to infer causal relationships using genetic variants as instrumental variables, yet the selection of genetic instruments is not always given sufficient attention. Many MR studies rely on default linkage disequilibrium (LD) clumping parameters (r2 <0.001, 10,000 kb), as implemented in commonly used tools, without assessment of their suitability for specific exposures. We investigated whether this approach yields optimal instruments or whether a more pragmatic strategy yields stronger instruments. Using UK Biobank data, we examined three distinct exposure types-circulating amino acids, body mass index (BMI), and major depressive disorder (MDD). For each phenotype, we systematically varied LD clumping thresholds (r2 and genomic distance) and evaluated each instrument via both their average strength (F-statistic) and total strength (R2). Across all phenotypes, optimal instruments differed from default parameters and varied by exposure. For amino acids and BMI, more stringent LD thresholds (r2=0.00001) combined with larger clumping windows improved instrument strength, whereas for MDD, a highly polygenic, binary trait, smaller windows with stringent r2 maximized variance explained while maintaining F-statistics above the desired threshold (>10). Notably, increasing the number of SNPs did not consistently improve instrument quality, highlighting a trade-off between instrument strength and potential pleiotropy. We demonstrate that universal reliance on default LD clumping parameters can lead to suboptimal instruments. We propose a pragmatic framework for instrument selection based on empirical evaluation of strength metrics, improving the robustness and transparency of MR analyses across different exposure types.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking Mendelian Randomization methods for causal inference using genome-wide association study summary statistics 93%
- A Scalable Framework for Identifying Allelic Series from Summary Statistics 93%
- The Causal Pivot: A Structural Approach to Genetic Heterogeneity and Variant Discovery in Complex Diseases 93%
Similar papers in this journal
- Rare variants association testing for a binary outcome when pooling individual level data from heterogeneous studies 94%
- Meta-MultiSKAT: Multiple phenotype meta-analysis for region-based association test 94%
- Taking population stratification into account by local permutations in rare-variant association studies on small samples 93%
Similar papers in this journal
- A Comprehensive Evaluation of Methods for Mendelian Randomization Using Realistic Simulations and an Analysis of 38 Biomarkers for Risk of Type-2 Diabetes 93%
- Bias in two-sample Mendelian randomization when using heritable covariable-adjusted summary associations 93%
- An empirical investigation into the impact of winner's curse on estimates from Mendelian randomization 91%
Similar papers in this journal
- Bayesian Variable Selection with a Pleiotropic Loss Function in Mendelian Randomization 93%
- Mendelian Randomization with longitudinal exposure data: simulation study and real data application 92%
- Two-phase sample selection strategies for design and analysis in post-genome wide association fine-mapping studies 92%
Similar papers in this journal
- Case-only analysis of gene-environment interactions using polygenic risk scores 91%
- A Hierarchical Approach Using Marginal Summary Statistics for Multiple Intermediates in a Mendelian Randomization or Transcriptome Analysis 91%
- Analyses using multiple imputation need to consider missing data in auxiliary variables 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.