Back

Causal inference for multiple risk factors and diseases from genomics data

Machnik, N.; Mahmoudi, M.; Kraetschmer, I.; Bauer, M. J.; Robinson, M. R.

2023-12-08 genetics
10.1101/2023.12.06.570392 bioRxiv
Show abstract

Statistical causal learning in genome-wide association studies (GWAS) relies on the instrumental variable method of Mendelian Randomization (MR). Currently, an over-whelming number of MR studies purport to show causal relationships among a wide range of risk factors and outcomes. Here, we find that naive application of many recently proposed MR approaches results in numerous null relationships being discovered as highly significant. We show that a well-controlled error rate can be achieved through a graphical inference approach which: (i) selects a set of genetic instrumental variables (IVs) from GWAS summary static controlling for LD, linkage and pleiotropy; (ii) accommodates rare variants and binary outcomes in a principled way; (iii) distinguishes direct from indirect risk factors in very high-dimensional data; and (iv) identifies potential unobserved latent confounding. Only 20 minutes of wall-clock compute time is required for our Causal Inference GWAS (CI-GWAS) approach to jointly analyze a set of 9 common health risk factors, four common complex metabolic disease outcomes and 8.4M genetic variants recorded for 458,747 individuals in the UK Biobank. Genome-wide, we find that very few genetic variants are suitable MR IVs, with only 696 variants remaining when analyzing all traits jointly. While we replicate almost all paths previously found by CAUSE between risk factors and outcomes, we show that only few of these reflect direct adjacencies and that many cannot be distinguished from unmeasured confounding within the UK Biobank data. Our results suggest that well-curated longitudinal records and family data are likely needed to overcome the mixtures of temporal precedence and reverse-causality in biobank data. Our approach provides a first-step toward robust principled screening for potential causal links to understand the underlying nature of phenotypic correlations in biobank data.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.