Causal considerations can determine the utility of machine learning assisted GWAS
Mukherjee, S.; McCaw, Z.; Amar, D.; Dey, R.; Soare, T.; Xu, K.; Somineni, H.; Insitro Research Team, ; Eriksson, N.; O'Dushlaine, C.
Show abstract
Machine Learning (ML) is increasingly employed to generate phenotypes for genetic discovery, either by imputing existing phenotypes into larger cohorts or by creating novel phenotypes. While these ML-derived phenotypes can significantly increase sample size, and thereby empower genetic discovery, they can also inflate the false discovery rate (FDR). Recent research has focused on developing estimators that leverage both true and machine-learned phenotypes to properly control the type-I error. Our work complements these efforts by exploring how the true positive rate (TPR) and FDR depend on the causal relationships among the inputs to the ML model, the true phenotypes, and the environment. Using a simulation-based framework, we study architectures in which the machine-learned proxy phenotype is derived from biomarkers (i.e. inputs) either causally upstream or downstream of the target phenotype. We show that no inflation of the false discovery rate occurs when the proxy phenotype is generated from upstream biomarkers, but that false discoveries can occur when the proxy phenotype is generated from downstream biomarkers. Next, we show that power to detect variants truly associated with the target phenotype depends on its heritability and correlation with the proxy phenotype. However, the source of the correlation is key to evaluating a proxy phenotype's utility for genetic discovery. We demonstrate that evaluating machine-learned proxy phenotypes using out-of-sample predictive performance (e.g. phenotypic correlation) provides a poor lens on utility. This is because overall predictive performance does not differentiate between genetic and environmental correlation. In addition to parsing these properties of machine-learned phenotypes via simulations, we further illustrate them using real-world data from the UK Biobank.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Beyond SNP Heritability: Polygenicity and Discoverability of Phenotypes Estimated with a Univariate Gaussian Mixture Model 96%
- Noise-augmented directional clustering of genetic association data identifies distinct mechanisms underlying obesity 96%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 96%
Similar papers in this journal
- The Causal Pivot: A Structural Approach to Genetic Heterogeneity and Variant Discovery in Complex Diseases 97%
- An allelic series rare variant association test for candidate gene discovery 97%
- Welch-weighted Egger regression reduces false positives due to correlated pleiotropy in Mendelian randomization 97%
Similar papers in this journal
- A Bayesian Approach to Correcting the Attenuation Bias of Regression Using Polygenic Risk Score 95%
- Characterization of direct and/or indirect genetic associations for multiple traits in longitudinal studies of disease progression 95%
- Hidden structure in polygenic scores and the challenge of disentangling ancestry interactions in admixed populations 95%
Similar papers in this journal
- Identifying causal genotype-phenotype relationships for population-sampled parent-child trios 96%
- Assumptions about frequency-dependent architectures of complex traits bias measures of functional enrichment 95%
- Statistics to prioritize rare variants in family-based sequencing studies with disease subtypes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.