snpXplorer: an interactive platform for haplotype-aware exploration and integrated annotation of GWAS data
Tesi, N.; Green, G. S.; Salazar, A.; van der Lee, S. J.; Hulsman, M.; Holstege, H.; Reinders, M.
Show abstract
BackgroundGenome-wide association studies (GWAS) have identified thousands of loci associated with complex traits and diseases, yet translating these signals into biological insight remains challenging. Most associated variants are non-coding and reside in linkage disequilibrium (LD) blocks, where multiple correlated variants jointly contribute to association signals. These clusters, or haplotypes, may capture shared regulatory and functional contexts. Interpreting GWAS signals thus requires approaches that integrate regulatory, functional, and cross-trait evidence, while preserving the broader haplotypic context of disease-associated loci. At the same time, the rapid growth of publicly available GWAS summary statistics has enabled large-scale cross-trait analyses, but also introduced redundancy across closely related phenotypes. Efficient interpretation of GWAS data therefore requires tools that integrate heterogeneous data sources while preserving genomic and biological contexts. ResultsWe present snpXplorer, an interactive web platform for haplotype-aware exploration and annotation of GWAS data. The platform incorporates >10,000 GWAS datasets from OpenGWAS and enables multi-scale analysis across variants, haplotypes, genes, and traits. Key features include (i) a haplotype-based representation of association signals derived from LD structure, (ii) a unified variant annotation framework integrating clinical annotations (ClinVar), allele frequencies (gnomAD), functional predictions (CADD, AlphaGenome), quantitative trait loci (GTEx), structural variation, and GWAS associations, and (iii) cross-trait exploration using semantic similarity-based clustering of phenotypes. Use cases centered on Alzheimers disease illustrate this utility: for example, at the TMEM106B locus, snpXplorer identified a haplotype linked to eleven distinct traits, revealing synergistic pleiotropy across neurological and behavioral phenotypes alongside antagonistic pleiotropy with height. ConclusionssnpXplorer allows users to browse, filter, and inspect variant-, haplotype-, gene- and trait-level evidence, lowering the barrier to biological interpretation of GWAS results. Compared with existing tools that focus on specific aspects of GWAS interpretation, the strength of snpXplorer is that it reduces the need for fragmented queries across databases.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Multi-Tissue Neocortical Transcriptome-Wide Associations Study Implicates 8 Genes Across 6 Genomic Loci in Alzheimer's Disease 94%
- scGRNom: a computational pipeline of integrative multi-omics analyses for predicting cell-type disease genes and regulatory networks 94%
- An atlas connecting shared genetic architecture of human diseases and molecular phenotypes provides insight into COVID-19 susceptibility 94%
Similar papers in this journal
- Predicting Disease-Specific Histone Modifications and Functional Effects of Non-coding Variants by Leveraging DNA Language Models 95%
- Accurate characterization of expanded tandem repeat length and sequence through whole genome long-read sequencing on PromethION. 94%
- STRling: a k-mer counting approach that detects short tandem repeat expansions at known and novel loci 94%
Similar papers in this journal
- SparkINFERNO: A scalable high-throughput pipeline for inferring molecular mechanisms of non-coding genetic variants 95%
- Summary statistics from large-scale gene-environment interaction studies for re-analysis and meta-analysis 94%
- Incorporating family disease history and controlling case-control imbalance for population based genetic association studies 93%
Similar papers in this journal
- Large meta-analysis of genome-wide association studies expands knowledge of the genetic etiology of Alzheimer’s disease and highlights potential translational opportunities 94%
- Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts 93%
- Set-based rare variant association tests for biobank scale sequencing data sets 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.