DeepCAST-GWAS: Improving the Discovery of Genetic Associations Using Deep Learning-Based Regulatory SNP Prioritization
Rabuzin, L.; Heep, K.; Sigfstead, S.; Boeva, V.
Show abstract
Genome-wide association studies (GWAS) have uncovered numerous variants linked to complex traits, yet power remains limited by the large multiple testing burden and the inclusion of many variants with minimal regulatory impact. We present Deep learning-based Chromatin Accessibility SNP Targeting for GWAS (DeepCAST-GWAS), a framework that integrates functional annotations derived from deep learning models to improve both the yield and the reliability of GWAS findings. DeepCAST-GWAS uses SNP Activity Difference (SAD) scores from in silico mutagenesis with the Enformer model to estimate the predicted effect of each variant on chromatin accessibility across tissues, allowing statistical testing to focus on variants with stronger regulatory evidence. Using conservative family-wise error rate (FWER) control, DeepCAST-FWER produces fewer associations than existing power-boosting approaches, but the associations it reports replicate in larger cohort GWAS at substantially higher rates. For applications where discovery count is more important, DeepCAST-sFDR increases the number of genome-wide significant findings above baseline GWAS by using the Enformer SAD scores for stratified False Discovery Rate (sFDR) control. DeepCAST-sFDR achieves performance comparable to the strongest competing method, while maintaining reliability on par with a standard GWAS. Subsampling analyses across a wide range of traits confirm these improvements in both sensitivity and replicability. DeepCAST-GWAS offers a principled way to incorporate sequence-based regulatory predictions into population-scale association testing, demonstrating that chromatin accessibility activity scores can improve the stability of GWAS discoveries. The framework is made available at https://github.com/BoevaLab/DeepCAST-GWAS.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SUMMIT: An integrative approach for better transcriptomic data imputation improves causal gene identification 96%
- XMAP: Cross-population fine-mapping by leveraging genetic diversity and accounting for confounding bias 96%
- Multi-context genetic modeling of transcriptional regulation resolves novel disease loci 96%
Similar papers in this journal
- Genotype inference from aggregated chromatin accessibility data reveals genetic regulatory mechanisms 96%
- Optimizing and benchmarking polygenic risk scores with GWAS summary statistics 95%
- Primo: integration of multiple GWAS and omics QTL summary statistics for elucidation of molecular mechanisms of trait-associated SNPs and detection of pleiotropy in complex traits 95%
Similar papers in this journal
- SparsePro: an efficient fine-mapping method integrating summary statistics and functional annotations 98%
- Enhancing Portability of Trans-Ancestral Polygenic Risk Scores through Tissue-Specific Functional Genomic Data Integration 96%
- Leveraging expression from multiple tissues using sparse canonical correlation analysis (sCCA) and aggregate tests improves the power of transcriptome-wide association studies (TWAS) 96%
Similar papers in this journal
- Fast and flexible joint fine-mapping of multiple traits via the Sum of Single Effects model 97%
- Mendelian randomization accounting for correlated and uncorrelated pleiotropic effects using genome-wide summary statistics. 96%
- MultiSuSiE improves multi-ancestry fine-mapping in All of Us whole-genome sequencing data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.