Quantile Regression for biomarkers in the UK Biobank
Wang, C.; Wang, T.; Wei, Y.; Aschard, H.; Ionita-Laza, I.
Show abstract
Genome-wide association studies (GWAS) for biomarkers important for clinical phenotypes can lead to clinically relevant discoveries. GWAS for quantitative traits are based on simplified regression models modeling the conditional mean of a phenotype as a linear function of genotype. An alternative and easy to apply approach is quantile regression that naturally extends linear regression to the analysis of the entire conditional distribution of a phenotype of interest by modeling conditional quantiles within a regression framework. Quantile regression can be applied efficiently at biobank scale using standard statistical packages in much the same way as linear regression, while having some unique advantages such as identifying variants with heterogeneous effects across different quantiles, including non-additive effects and variants involved in gene-environment interactions; accommodating a wide range of phenotype distributions with invariance to trait transformation; and overall providing more detailed information about the underlying genotype-phenotype associations. Here, we demonstrate the value of quantile regression in the context of GWAS by applying it to 39 quantitative traits in the UK Biobank (n > 300, 000 individuals). Across these 39 traits we identify 7,297 significant loci, including 259 loci only detected by quantile regression. We show that quantile regression can help uncover replicable but unmodelled gene-environment interactions, and can provide additional key insights into poorly understood genotype-phenotype correlations for clinically relevant biomarkers at minimal additional cost.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 98%
- OTTERS: A powerful TWAS framework leveraging summary-level reference data 97%
- A novel Mendelian randomization method identifies causal relationships between gene expression and low-density lipoprotein cholesterol levels. 97%
Similar papers in this journal
Similar papers in this journal
- Valid inference for machine learning-assisted GWAS 98%
- Leveraging functional genomic annotations and genome coverage to improve polygenic prediction of complex traits within and between ancestries 97%
- A new method for multi-ancestry polygenic prediction improves performance across diverse populations 97%
Similar papers in this journal
- Polygenic scores capture genetic modification of the adiposity-cardiometabolic risk factor relationship 97%
- Integrative polygenic risk score improves the prediction accuracy of complex traits and diseases 96%
- MUSSEL: Enhanced Bayesian Polygenic Risk Prediction Leveraging Information across Multiple Ancestry Groups 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.