Improving the discovery of rare variants associated with alcohol problems by leveraging machine learning phenotype prediction and functional information.
Ahangari, M.; Gentry, A. E.; Hassan, M.; Nguyen, T. H.; Kendler, K. S.; Bacanu, S.-A.; Peterson, R. E.; Riley, B. P.; Webb, B. T.
Show abstract
Alcohol use disorder (AUD) is moderately heritable with significant social and economic impact. Genome-wide association studies (GWAS) have identified common variants associated with AUD, however, rare variant investigations have yet to achieve well-powered sample sizes. In this study, we conducted an interval-based exome-wide analysis of the Alcohol Use Disorder Identification Test Problems subscale (AUDIT-P) using both machine learning (ML) predicted risk and empirical functional weights. This research has been conducted using the UK Biobank Resource (application number 30782.) Filtering the 200k exome release to unrelated individuals of European ancestry resulted in a sample of 147,386 individuals with 51,357 observed and 96,029 unmeasured but predicted AUDIT-P for exome analysis. Sequence Kernel Association Test (SKAT/SKAT-O) was used for rare variant (Minor Allele Frequency (MAF) < 0.01) interval analyses using default and empirical weights. Empirical weights were constructed using annotations found significant by stratified LD Score Regression analysis of predicted AUDIT-P GWAS, providing prior functional weights specific to AUDIT-P. Using only samples with observed AUDIT-P yielded no significantly associated intervals. In contrast, ADH1C and THRA gene intervals were significant (False discovery rate (FDR) <0.05) using default and empirical weights in the predicted AUDIT-P sample, with the most significant association found using predicted AUDIT-P and empirical weights in the ADH1C gene (SKAT-O P Default= 1.06 x 10-9 and P Empirical weight = 6.25 x 10-11). These findings provide evidence for rare variant association of the ADH1C gene with the AUDIT-P and highlight the successful leveraging of ML to increase effective sample size and prior empirical functional weights based on common variant GWAS data to refine and increase the statistical significance in underpowered phenotypes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genome-wide association study of problematic opioid prescription use in 132,113 23andMe research participants of European ancestry 95%
- Intergenerational transmission of complex traits and the offspring methylome 94%
- Genome-wide association study and multi-trait analysis of opioid use disorder identifies novel associations in 639,709 individuals of European and African ancestry 93%
Similar papers in this journal
- Multi-trait genome-wide association analyses leveraging alcohol use disorder findings identify novel loci for smoking behaviors in the Million Veteran Program 95%
- Gene-based polygenic risk scores analysis of alcohol use disorder in African Americans 94%
- Divergent gene expression in alcohol and opioid usedisorders results in consistent alterations in functional networks in the Dorsolateral Prefrontal Cortex 93%
Similar papers in this journal
- Functional classes of SNPs related to psychiatric disorders and behavioral traits contrast with those related to neurological disorders 93%
- Assessing the performance of genome-wide association studies for predicting disease risk 93%
- Alcohol use and cardiometabolic risk in the UK Biobank: a Mendelian randomization study 92%
Similar papers in this journal
- "The Heidelberg Five" Personality Dimensions: Genome-wide Associations, Polygenic Risk for Neuroticism, and Psychopathology 20 Years after Assessment 94%
- Schizophrenia Risk Alleles Often Affect The Expression of Many Genes and Each Gene May Have a Different Effect On The Risk; A Mediation Analysis. 93%
- Improving Machine Learning Prediction of ADHD Using Gene Set Polygenic Risk Scores and Risk Scores from Genetically Correlated Phenotypes 93%
Similar papers in this journal
- Methylome-wide association study of early life stressors and adult mental health reveals a relationship between birth date and cell type composition in blood 93%
- Imputed Gene Expression Risk Scores: A Functionally Informed Component of Polygenic Risk 93%
- Common genetic variants with fetal effects on birth weight are enriched for proximity to genes implicated in rare developmental disorders 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.