Statistical model integrating interactions into genotype-phenotype association mapping: an application to reveal 3D-genetic basis underlying Autism
Li, Q.; Cao, C.; Perera, D.; He, J.; Chen, X.; Azeem, F.; Howe, A.; Au, B.; Yan, J.; Long, Q.
Show abstract
Biological interactions are prevalent in the functioning organisms. Correspondingly, statistical geneticists developed various models to identify genetic interactions through genotype-phenotype association mapping. The current standard protocols in practice test single variants or single regions (that contain multiple local variants) sequentially along the genome, followed by functional annotations that involve various aspects including interactions. The testing of genetic interactions upfront is rare in practice due to the burden of testing a huge number of combinations, which lead to the multiple-test problem and the risk of overfitting. In this work, we developed interaction-integrated linear mixed model (ILMM), a novel model that integrates a priori knowledge into linear mixed models. ILMM enables statistical integration of genetic interactions upfront and overcomes the problems associated with combination searching. Three dimensional (3D) genomic interactions assessed by Hi-C experiments have led to unprecedented biological discoveries. However, the contribution of 3D genomic interactions to the genetic basis of complex diseases has yet to be quantified. Using 3D interacting regions as a priori information, we conducted both simulations and real data analysis to test ILMM. By applying ILMM to whole genome sequencing data for Autism Spectrum Disorders, or ASD (MSSNG) and transcriptome sequencing data (GTEx), we revealed the 3D-genetic basis of ASD and 3D-eQTLs for a substantial proportion of gene expression in brain tissues. Moreover, we have revealed a potential mechanism involving distal regulation between FOXP2 and DNMT3A conferring the risk of ASD. Software is freely available in our GitHub: https://github.com/theLongLab/Jawamix5
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A machine learning approach to predicting autism risk genes: Validation of known genes and discovery of new candidates 92%
- MetaPhat: Detecting and decomposing multivariate associations from univariate genome-wide association statistics 92%
- Integrative Ranking Of Enhancer Networks Facilitates The Discovery Of Epigenetic Markers In Cancer 92%
Similar papers in this journal
- Bayesian estimation of cell-type-specific gene expression per bulk sample with prior derived from single-cell data 96%
- Co-expression enrichment analysis at the single-cell level reveals convergent defects in neural progenitor cells and their cell-type transitions in neurodevelopmental disorders 94%
- An Association Test of the Spatial Distribution of Rare Missense Variants within Protein Structures Improves Statistical Power of Sequencing Studies 93%
Similar papers in this journal
Similar papers in this journal
- CausalCell: applying causal discovery to single-cell analyses 93%
- MCGA: a multi-strategy conditional gene-based association framework integrating with isoform-level expression profiles reveals new susceptible and druggable candidate genes of schizophrenia 93%
- Novel genetic loci affecting facial shape variation in humans 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.