A Semi-Supervised Ensemble Approach to Rank Potential Causal Variants and Their Target Genes in Microglia for Alzheimer's Disease
Khaire, A.; Wen, J.; Yang, X.; Shen, Y.; Li, Y.
Show abstract
Alzheimers disease (AD) is the leading cause of death among individuals over 65. Despite many AD genetic variants detected by large genome-wide association studies (GWAS), a limited number of causal genes have been confirmed. Conventional machine learning techniques integrate functional annotation data and GWAS signals to assign variants functional relevance probabilities. Yet, a large proportion of genetic variation lies in the non-coding genome, where unsupervised and semi-supervised techniques have demonstrated greater advantage. Furthermore, cell-type specific approaches are needed to better understand disease etiology. Studying AD from a microglia-specific lens is more likely to reveal causal variants involved in immune pathways. Therefore, in this study, we developed S-BEAM: a semi-supervised ensemble approach using microglia-specific data to prioritize non-coding variants and their target genes that play roles in immune-related AD mechanisms. We designed a transductive positive-unlabeled and negative-unlabeled learning model that employs a bagging technique to learn from unlabeled variants, generating multiple predicted probabilities of variant risk. Using a combined homogeneous-heterogeneous ensemble framework, we aggregated the predictions. We applied our model to AD variant data, identifying 11 risk variants acting in well-known AD genes, such as TSPAN14, INPP5D, and MS4A2. These results validated our models performance and demonstrated a need to study these genes in the context of microglial pathways. We also proposed further experimental study for 37 potential causal variants associated with less-known genes. Our work has utility in predicting AD relevant genes and variants functioning in microglia and can be generalized for application to other complex diseases or cell types.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Module analysis using single-patient differential expression signatures improve the power of association study for Alzheimer's disease 93%
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 91%
- Hardy-Weinberg Equilibrium in the Large Scale Genomic Sequencing Era 90%
Similar papers in this journal
- DYNATE: Localizing Rare-Variant Association Regions via Multiple Testing Embedded in an Aggregation Tree 92%
- Polygenic hazard score models for the prediction of Alzheimer's free survival using the lasso for Cox's proportional hazards model 91%
- RetroFun-RVS: a retrospective family-based framework for rare-variant analysis incorporating functional annotations 90%
Similar papers in this journal
- Identifying and ranking potential driver genes of Alzheimer's Disease using multi-view evidence aggregation 94%
- CLEP: A Hybrid Data- and Knowledge- Driven Framework for Generating Patient Representations 94%
- Deep5hmC: Predicting genome-wide 5-Hydroxymethylcytosine landscape via a multimodal deep learning model 93%
Similar papers in this journal
- Artificial intelligence-driven meta-analysis of brain gene expression data identifies novel gene candidates in Alzheimer’s Disease 94%
- Identifying Alzheimer's disease-related pathways based on whole-genome sequencing data 94%
- Multi-task deep autoencoder to predict Alzheimer’s disease progression using temporal DNA methylation data in peripheral blood 92%
Similar papers in this journal
- Genotyping TOMM40’523 Poly-T Polymorphisms Using Whole-Genome Sequencing 95%
- Integrating spatial transcriptomics and snRNA-seq data enhances differential gene expression analysis results of AD-related phenotypes 94%
- A Specialized Reference Panel with Structural Variants Integration for Improving Genotype Imputation in Alzheimer's Disease and Related Dementias (ADRD) 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.