Hi-GREx: A 3D Genome Guided Framework for enhancing Gene Expression Prediction Using Hi-C Selected Distal SNPs
Joshi, K.; Xuan, Z.; Chen, M.
Show abstract
Genome-Wide Association Study (GWAS) method has been successfully used to map thousands of loci associated with complex traits, but its ability to reveal the molecular mechanisms altered in complex diseases has been limited due to not including combinations and interactions between markers when predicting a disease. Transcriptome-Wide Association Studies (TWAS) estimate the aggregate effects of multiple genetic variants on complex diseases and represent a promising approach to address the limitations of GWAS. In particular, TWAS provides insights into the functional consequences of disease-associated SNPs by linking them to gene transcription, thereby offering a mechanistic understanding that GWAS alone cannot provide. However, TWAS associated variants have been annotated with the closest or most biologically relevant candidate gene within arbitrarily defined distances but fails to account for long distance SNPs which can affect many genes and have a widespread impact on regulatory networks. Therefore, there is a need to leverage these observed enrichments and build a method that incorporates both short and long distance-associations between SNPs and complex phenotypes. Here we present a method which can utilize Hi-C data to capture "informative" long-distance SNPs and aim to improve prediction accuracy of previous TWAS method. We benchmarked our method on GTEx brain cortex genotype and expression data together with the corresponding Hi-C data. By using the "informative" long distance SNPs selected based on Hi-C, our method improved prediction accuracy of gene expression for 77.4% of the active genes across the entire genome. Particularly, our method can build significant expression models for 18% of genes which were missed by using only short-distance SNPs. Our method has demonstrated the efficiency and importance of utilizing long-distance SNPs in predicting gene expression and can further enhance the power of TWAS methods.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Predicting Disease-Specific Histone Modifications and Functional Effects of Non-coding Variants by Leveraging DNA Language Models 97%
- HyperChIP for identifying hypervariable signals across ChIP/ATAC-seq samples 95%
- Genotype inference from aggregated chromatin accessibility data reveals genetic regulatory mechanisms 95%
Similar papers in this journal
- Co-expression-wide association studies link genetically regulated interactions with complex traits 96%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 95%
- SUMMIT: An integrative approach for better transcriptomic data imputation improves causal gene identification 95%
Similar papers in this journal
- Bayesian Estimation of Allele-Specific Expression in the Presence of Phasing Uncertainty 95%
- DeepPerVar: a multimodal deep learning framework for functional interpretation of genetic variants in personal genome 95%
- Deep5hmC: Predicting genome-wide 5-Hydroxymethylcytosine landscape via a multimodal deep learning model 95%
Similar papers in this journal
- Disease-specific prioritization of non-coding GWAS variants based on chromatin accessibility 94%
- Extensive co-regulation of neighbouring genes complicates the use of eQTLs in target gene prioritisation 94%
- Scalable Bayesian functional GWAS method accounting for multivariate quantitative functional annotations with applications to studying Alzheimer’s disease 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.