Disease association with frequented regions of genotype graphs
Hokin, S.; Cleary, A.; Mudge, J.
Show abstract
Complex diseases, with many associated genetic and environmental factors, are a challenging target for genomic risk assessment. Genome-wide association studies (GWAS) associate disease status with, and compute risk from, individual common variants, which can be problematic for diseases with many interacting or rare variants. In addition, GWAS typically employ a reference genome which is not built from the subjects of the study, whose genetic background may differ from the reference and whose genetic characterization may be limited. We present a complementary method based on disease association with collections of genotypes, called frequented regions, on a pangenomic graph built from subjects genomes. We introduce the pangenomic genotype graph, which is better suited than sequence graphs to human disease studies. Our method draws out collections of features, across multiple genomic segments, which are associated with disease status. We show that the frequented regions method consistently improves machine-learning classification of disease status over GWAS classification, allowing incorporation of rare or interacting variants. Notably, genomic segments that have few or no variants of genome-wide signif-icance (p < 5 x 10-8) provide much-improved classification with frequented regions, encouraging their application across the entire genome. Frequented regions may also be utilized for purposes such as choice of treatment in addition to prediction of disease risk.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Discovering genetic interactions bridging pathways in genome-wide association studies 97%
- Sharing information between related diseases using Bayesian joint fine mapping increases accuracy and identifies novel associations in six immune mediated diseases 95%
- MutPred2: inferring the molecular and phenotypic impact of amino acid variants 95%
Similar papers in this journal
- COBT: A gene-based rare variant burden test for case-only study designs using aggregated genotypes from public reference cohorts. 93%
- ClinSV: Clinical grade structural and copy number variant detection from whole genome sequencing data 93%
- An atlas connecting shared genetic architecture of human diseases and molecular phenotypes provides insight into COVID-19 susceptibility 93%
Similar papers in this journal
- GA4GH Phenopacket-Driven Characterization of Genotype-Phenotype Correlations in Mendelian Disorders 95%
- The Causal Pivot: A Structural Approach to Genetic Heterogeneity and Variant Discovery in Complex Diseases 94%
- Focus on single gene effects limits discovery and interpretation of complex trait-associated variants 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.