Clustering Of Rare Variants For Causal Variants Identification And Effect Direction Classification
Sun, X.; Liu, X.; Liu, C.
Show abstract
Several gene-based tests, e.g., sequence kernel association test, have been developed for association testing of rare single nucleotide variants (SNVs) in genomic regions with disease traits. A common limitation of these aggregate methods is their inability to discriminate potentially causal variants from null variants within the tested regions. We propose a novel clustering method to classify rare variants into null and signal variant groups using summary statistics from the gene-based tests based on a Gaussian mixture model (GMM). We classify the signal variants into potentially risk and protective subgroups of different effect sizes. We evaluate the performance of the proposed method by a simulation study, considering several statistics such as the adjusted rand index (ARI), mean square error (MSE), and accuracy in specifying the number of clusters. We apply the proposed clustering method to identify possibly risk and protective rare variants in six genes that are significantly associated with blood pressure (BP) traits in the most recent large genomewide association study (GWAS) and meta-analysis. This proposed method may facilitate the identification of potentially causal rare variant clusters in genomic regions and ultimately help understand the genetic architecture underlying human complex traits for the discovery of drug target and the design of gene therapy.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Noise-augmented directional clustering of genetic association data identifies distinct mechanisms underlying obesity 95%
- Joint Modeling of Effect Sizes for Two Correlated Traits: Characterizing Trait Properties to Enhance Polygenic Risk Prediction 94%
- A novel method for multiple phenotype association studies based on genotype and phenotype network 94%
Similar papers in this journal
- Identity-by-descent mapping using multi-individual IBD with genome-wide multiple testing adjustment 95%
- GxE PRS: Genotype-environment interaction in polygenic risk score models for quantitative and binary traits 95%
- Quantifying posterior effect size distribution of susceptibility loci by common summary statistics 95%
Similar papers in this journal
- BinomiRare: A carriers-only test for association of rare genetic variants with a binary outcome for mixed models and any case-control proportion 94%
- A parametric bootstrap approach for computing confidence intervals for genetic correlations with application to genetically-determined protein-protein networks 94%
- Leveraging TOPMed Imputation Server and Constructing a Cohort-Specific Imputation Reference Panel to Enhance Genotype Imputation among Cystic Fibrosis Patients 93%
Similar papers in this journal
- Exploiting Family History in Aggregation Unit-based Genetic Association Tests 95%
- Lifestyle Risk Score for aggregating multiple lifestyle factors: Handling missingness of individual lifestyle components in meta-analysis of gene-by-lifestyle interactions 93%
- Prioritization of disease genes from GWAS using ensemble based positive-unlabeled learning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.