Leveraging functional annotations to map rare variants associated with Alzheimer's disease with gruyere
Das, A.; Lakhani, C.; Terwagne, C.; Lin, J.-S.; Naito, T.; Raj, T.; Knowles, D. A.
Show abstract
The increasing availability of whole-genome sequencing (WGS) has begun to elucidate the contribution of rare variants (RVs), both coding and non-coding, to complex disease. Multiple RV association tests are available to study the relationship between genotype and phenotype, but most are restricted to per-gene models and do not fully leverage the availability of variant-level functional annotations. We propose Genome-wide Rare Variant EnRichment Evaluation (gruyere), a Bayesian probabilistic model that complements existing methods by learning global, trait-specific weights for functional annotations to improve variant prioritization. We apply gruyere to WGS data from the Alzheimers Disease (AD) Sequencing Project, consisting of 7,966 cases and 13,412 controls, to identify AD-associated genes and annotations. Growing evidence suggests that disruption of microglial regulation is a key contributor to AD risk, yet existing methods have not had sufficient power to examine rare non-coding effects that incorporate such cell-type specific information. To address this gap, we 1) use predicted enhancer and promoter regions in microglia and other potentially relevant cell types (oligodendrocytes, astrocytes, and neurons) to define per-gene non-coding RV test sets and 2) include cell-type specific variant effect predictions (VEPs) as functional annotations. gruyere identifies 15 significant genetic associations not detected by other RV methods and finds deep learning-based VEPs for splicing, transcription factor binding, and chromatin state are highly predictive of functional non-coding RVs. Our study establishes a novel and robust framework incorporating functional annotations, coding RVs, and cell-type associated non-coding RVs, to perform genome-wide association tests, uncovering AD-relevant genes and annotations.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting Disease-Specific Histone Modifications and Functional Effects of Non-coding Variants by Leveraging DNA Language Models 96%
- Cross-species imputation and comparison of single-cell transcriptomic profiles 95%
- INFIMA leverages multi-omics model organism data to identify effector genes of human GWAS variants 94%
Similar papers in this journal
- Unsupervised representation learning improves genomic discovery and risk prediction for respiratory and circulatory functions and diseases 94%
- Fast and flexible joint fine-mapping of multiple traits via the Sum of Single Effects model 94%
- Personal transcriptome variation is poorly explained by current genomic deep learning models 93%
Similar papers in this journal
Similar papers in this journal
- Redefining tissue specificity of genetic regulation of gene expression in the presence of allelic heterogeneity 95%
- A unified framework for cell-type-specific eQTLs prioritization by integrating bulk and scRNA-seq data 94%
- Bayesian transcriptome-wide association study method leveraging both cis- and trans- eQTL information through summary statistics 93%
Similar papers in this journal
- scGRNom: a computational pipeline of integrative multi-omics analyses for predicting cell-type disease genes and regulatory networks 95%
- Multi-Tissue Neocortical Transcriptome-Wide Associations Study Implicates 8 Genes Across 6 Genomic Loci in Alzheimer's Disease 95%
- Leveraging genomic diversity for discovery in an EHR-linked biobank: the UCLA ATLAS Community Health Initiative 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.