Machine Learning-Based Prediction of Cell-type Resolved Brain eQTLs Enhances Discovery of Variants Explaining Alzheimer's Disease Heritability
Lakhani, C. M.; Cavalca, G.; Liu, A.; Nidumbur, R.; Feng, R.; Raj, T.; De Jager, P.; The Alzheimer's Disease Functional Genomics Consortium, ; Wang, G.; Knowles, D. A.
Show abstract
The majority of causal genome-wide association studies (GWAS) variants for Alzheimers disease (AD) are believed to reside in noncoding regions of the genome, where they likely affect gene regulation, particularly in microglia. Although expression Quantitative Trait Loci (eQTL) studies offer valuable insights into gene regulation, they tend to identify variants in the promoter regions of genes under weaker selection. In contrast, GWAS variants are often found in enhancer regions linked to genes under stronger selection. To address this discrepancy, we developed predictive models, called single-cell Enhanced Expression Modifier Scores (scEEMS), to identify cell type-specific eQTLs using 4,839 genomic features, including deep learning-based scores that predict the effects of variants on various molecular phenotypes. These models were trained on fine-mapped single-cell eQTLs from six cell types and exhibited strong performance, with an average cross-validation area under the precision-recall curve (AUPRC) of 0.67. Notably, for microglia, the predicted eQTLs explained 15.3% of AD GWAS heritability, with a 140.6-fold enrichment of heritability (p-value 5.15 x 10-6), compared to just 5% of heritability explained by fine-mapped eQTLs. Incorporating scEEMS predictions as priors in eQTL fine-mapping refined credible sets and resulted in a net gain of 107 eGenes in microglia and 271 eGenes in astrocytes, reflecting improved statistical power that both identified new associations and filtered false positives. We used these models to link variants to cell type-specific genes and then applied the eMAGMA framework to nominate cell type-specific AD risk genes based on GWAS data. Our eMAGMA approach using predicted eQTLs identified 215 cell type-gene pairs (111 unique genes), substantially more than the 76 pairs (55 unique genes) found using fine-mapped eQTLs. Among these, we identified 43 microglia-specific, 18 astrocyte-specific, and 15 oligodendrocyte-specific AD risk genes. Of the 215 pairs, 62 replicated in at least one non-European population (African-American, Hispanic, or East Asian), compared to only 29 from fine-mapped eQTLs. Importantly, 18 of the replicated pairs represent novel discoveries not found by standard MAGMA or fine-mapped eMAGMA analyses. Five cell type-gene pairs replicated across European and at least two non-European populations: BIN1 (microglia), PICALM (microglia), ABCA7 (astrocyte and oligodendrocyte), and ARHGAP45 (microglia).
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genetics of the human microglia regulome refines Alzheimer’s disease risk loci 98%
- Atlas of genetic effects in human microglia transcriptome across brain regions, aging and disease pathologies 98%
- Spatial and single-nucleus transcriptomic analysis of genetic and sporadic forms of Alzheimer's Disease 97%
Similar papers in this journal
Similar papers in this journal
- Interaction molecular QTL mapping discovers cellular and environmental modifiers of genetic regulatory effects 96%
- A Transcriptomic Atlas of the Human Brain Reveals Genetically Determined Aspects of Neuropsychiatric Health 96%
- Brain eQTLs of European, African American, and Asian ancestry improve interpretation of schizophrenia GWAS 95%
Similar papers in this journal
- Meta-analysis fine-mapping is often miscalibrated at single-variant resolution 96%
- Variant-resolved prediction of context-specific isoform variation with a graph-based attention model 96%
- Polygenic regression uncovers trait-relevant cellular contexts through pathway activation transformation of single-cell RNA sequencing data 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.