Separating the genetics of disease, treatment and treatment response using graphical modeling and large-scale electronic health records.
Borczyk, M.; Machnik, N.; Hajto, J.; Kraetschmer, I.; Konowalska, P.; Baszkiewicz, B.; Korostynski, M.; Robinson, M. R.
Show abstract
Genetic variants affect baseline health and biomarker values, which in turn may impact both the therapy selected for an individual and the magnitude of change induced by the medication. Here, we propose an approach for complex longitudinal repeated measures biobank data, which separates genetic effects for disease from the genetic effects for medication usage and those for treatment response. For 211,845 individuals, we construct a pre-post study design from 1,420,443 repeated blood pressure (BP) measurements and 1,117,900 prescription records for common BP influencing drugs, using electronic health records. We model these jointly alongside 8,430,446 imputed single nucleotide polymorphism (SNP) markers and 17,852 whole-exome sequence loss-of-function (LoF) variants, all within a single novel graphical modeling framework. We identify pharmacogenetic candidate SNPs and LoF variants in genes SLC35F2, PKD1 and KCNIP4, which are associated with angiotensin receptor blocker therapy and response after controlling for hypertensive disease status across multiple world-wide biobanks. We additionally detect and replicate established clinically relevant variants for statin treatment across multiple biobanks. We find that genetic variation for BP is predominantly shaped prior to the age of 50, but we identify 127 independent loci associated with age-specific BP changes later in life. Finally, once post-treatment measures are conditioned on pre-treatment measures and therapy, we find evidence for four independent loci influencing BP treatment response, including a variant in ADAMTSL1 which has previously been associated with diuretic and beta-blocker response. Our graphical modeling and pre-post study design provides a robust way of detecting time-, treatment- and treatment response-specific genetic associations within large-scale biobank studies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Unsupervised representation learning improves genomic discovery and risk prediction for respiratory and circulatory functions and diseases 96%
- Combining case-control status and family history of disease increases association power 95%
- Mendelian randomization accounting for correlated and uncorrelated pleiotropic effects using genome-wide summary statistics. 95%
Similar papers in this journal
- Co-expression-wide association studies link genetically regulated interactions with complex traits 95%
- Testing and controlling for horizontal pleiotropy with the probabilistic Mendelian randomization in transcriptome-wide association studies 95%
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 95%
Similar papers in this journal
Similar papers in this journal
- Leveraging genomic diversity for discovery in an EHR-linked biobank: the UCLA ATLAS Community Health Initiative 94%
- An atlas connecting shared genetic architecture of human diseases and molecular phenotypes provides insight into COVID-19 susceptibility 93%
- A loss-of-function CCR2 variant is associated with lower cardiovascular risk 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.